Artificial intelligence-based research and development data management method

By processing structured and unstructured R&D data in a data middle platform and data lake warehouse integrated platform, generating analyzable datasets, and using an artificial intelligence decision engine to execute intelligent processing processes, the problem of unified data management is solved, and data accumulation and optimization are achieved.

CN122132600APending Publication Date: 2026-06-02SUZHOU CHIMA TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SUZHOU CHIMA TECHNOLOGY CO LTD
Filing Date
2026-03-02
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

In R&D management scenarios, the differences between structured and unstructured R&D data make it difficult to manage data in a unified manner, making it impossible to form a dataset that can be used for analysis and modeling, and making it difficult to accumulate and optimize management results.

Method used

Input data is accessed through a data acquisition bus and then stored and managed in the integrated data lake warehouse platform. Data cleaning, deduplication, standardization, tagging, and versioning are performed to generate R&D datasets that can be used for analysis and modeling. The AI ​​decision engine is then used to execute feature processing and intelligent processing procedures, and the results are finally written back to the data platform.

Benefits of technology

It enables centralized accumulation and reuse of R&D data, provides a basis for sustainable dynamic optimization, and supports the dynamic optimization of subsequent R&D data management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122132600A_ABST
    Figure CN122132600A_ABST
Patent Text Reader

Abstract

The application provides a research and development data management method based on artificial intelligence. The method obtains input data of a research and development management object, accesses the input data to a data hub through a data collection bus, stores and manages the input data in a data lake warehouse integrated platform, obtains a research and development data set, performs feature processing based on the research and development data set, obtains feature data, inputs the feature data into an artificial intelligence decision engine to execute at least one intelligent processing process, outputs a management result corresponding to the intelligent processing process, and writes the management result back to the data hub, so that the centralized precipitation and reusability of the management result can be realized, a sustainable dynamic optimization basis for subsequent research and development data management is provided, and the subsequent research and development data management can be dynamically optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to data processing technology, and more particularly to a research and development data management method based on artificial intelligence. Background Technology

[0002] In R&D management scenarios, the data corresponding to R&D management objects usually includes both structured and / or unstructured R&D data, which come from different business processes and information systems.

[0003] Structured and unstructured R&D data differ in their data organization, semantic granularity, and computability, making it difficult to directly form a unified, manageable, and analytically applicable R&D dataset. Furthermore, without a unified R&D dataset, input data is difficult to transform into usable feature data for intelligent processing when performing feature processing and input mechanisms for AI decision engines. Management results are also difficult to write back and consolidate, hindering the formation of a data loop for subsequent dynamic optimization. Summary of the Invention

[0004] This application provides an artificial intelligence-based R&D data management method to achieve centralized accumulation and reusability of management results, thereby providing a sustainable and dynamic optimization basis for subsequent R&D data management, and thus enabling dynamic optimization of subsequent R&D data management.

[0005] Firstly, this application provides a research and development data management method based on artificial intelligence, including: Obtain input data from the R&D management object, wherein the input data includes at least R&D data; The input data is connected to the data platform via a data acquisition bus and stored and managed in the integrated data lake warehouse platform to obtain the R&D dataset. Feature data is obtained by performing feature processing on the aforementioned R&D dataset. The feature data is input into an artificial intelligence decision engine to execute at least one intelligent processing procedure; Output the management results corresponding to the intelligent processing flow, and write the management results back to the data platform.

[0006] Secondly, this application provides an artificial intelligence-based R&D data management device, comprising: The acquisition module is used to acquire input data of the R&D management object, wherein the input data includes at least R&D data; The processing module is used to connect the input data to the data platform through the data acquisition bus, and store and manage it in the integrated data lake warehouse platform to obtain the R&D dataset; The processing module is also used to perform feature processing based on the R&D dataset to obtain feature data; An execution module is used to input the feature data into an artificial intelligence decision engine to execute at least one intelligent processing flow; The output module is used to output the management results corresponding to the intelligent processing flow and write the management results back to the data platform.

[0007] Thirdly, this application provides an electronic device, comprising: Processor; and, Memory for storing the executable instructions of the processor; The processor is configured to perform any of the possible methods described in the first aspect by executing the executable instructions.

[0008] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement any of the possible methods described in the first aspect.

[0009] The AI-based R&D data management method provided in this application acquires input data from R&D management objects, connects the input data to a data platform via a data acquisition bus, and stores and manages it in a data lake warehouse integrated platform to obtain an R&D dataset. Feature processing is then performed on the R&D dataset to obtain feature data. This feature data is then input into an AI decision engine to execute at least one intelligent processing flow, outputting management results corresponding to the intelligent processing flow. The management results are then written back to the data platform, enabling centralized accumulation and reusability of management results. This provides a sustainable and dynamic optimization basis for subsequent R&D data management, allowing for dynamic optimization of subsequent R&D data management. Attached Figure Description

[0010] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0011] Figure 1 This application illustrates a flowchart of an artificial intelligence-based R&D data management method according to an example embodiment. Figure 2 This is a flowchart illustrating the implementation of S140 according to the first example embodiment of this application; Figure 3 This is a flowchart illustrating the implementation of S140 according to the second example embodiment of this application; Figure 4This is a flowchart illustrating the implementation of S140 according to the third exemplary embodiment of this application; Figure 5 This is a flowchart illustrating the implementation of S140 according to the fourth exemplary embodiment of this application; Figure 6 This is a flowchart illustrating the implementation of S140 according to the fifth exemplary embodiment of this application; Figure 7 This is a flowchart illustrating the implementation of S140 according to the sixth exemplary embodiment of this application; Figure 8 This is a schematic diagram of the structure of an artificial intelligence-based R&D data management device according to an example embodiment of this application; Figure 9 This is a schematic diagram of the structure of an electronic device according to an example embodiment of this application.

[0012] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0013] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0014] Figure 1 This is a flowchart illustrating an artificial intelligence-based R&D data management method according to an example embodiment of this application. Figure 1 As shown, the method provided in this embodiment includes: S110. Obtain input data for R&D management objects.

[0015] In this step, the input data of the R&D management object may be obtained, wherein the input data includes at least structured R&D data and / or unstructured R&D data.

[0016] Specifically, this could involve receiving user-submitted text requirements and / or voice commands and / or project process data through a user interaction layer. For voice commands, this could involve speech-to-text processing to obtain the requirement text. Then, the requirement text and project process data could be combined to form unstructured R&D data and / or structured R&D data.

[0017] S120. Input data is connected to the data platform via the data acquisition bus and stored and managed in the integrated data lake warehouse platform to obtain the R&D dataset.

[0018] In this step, input data can be connected to the data platform via a data acquisition bus and stored and managed in the integrated data lake warehouse platform to obtain R&D datasets that can be used for analysis and modeling. Specifically, this can involve data cleaning, deduplication, standardization, and access control processing of the input data. Then, based on a pre-defined data model, the input data is labeled and versioned to form a research and development dataset.

[0019] Optionally, the above data cleaning of the input data includes: Perform field integrity and format validation on the input data; The system identifies and processes null values, outliers, and noise values ​​in the input data, and performs missing value filling, outlier correction, and / or noise filtering based on preset cleaning rules. The input data is timestamped and encoded to form cleaned input data.

[0020] Optionally, the above deduplication of input data includes: Generate deduplication features for data records in the input data. The deduplication features include at least one or more of the following: primary key identifier, content hash value, and time window identifier. Perform exact matching deduplication based on deduplication features and / or fuzzy matching deduplication based on similarity thresholds; The system retains the target data record for deduplication conflicting data records according to a preset retention strategy, and outputs the deduplicated input data.

[0021] Optionally, the above standardization of input data includes: Input data is mapped to a preset data model based on a preset field mapping table; Perform data type standardization, unit of measurement standardization, and enumeration value standardization on the input data; Based on the preset coding standard, personnel identifiers, task identifiers, project identifiers and document identifiers are uniformly coded to output standardized input data.

[0022] Optionally, the above-mentioned access control processing for input data includes: Configure data access policies for input data that correspond to user roles. Data access policies should include at least read permissions, write permissions, and export permissions. Perform row-level and / or column-level access control on the input data based on the data access policy; Audit logs are generated for access operations to input data and written to the audit log database.

[0023] Optionally, the above-mentioned labeling of input data based on a preset data model includes: Data tags are generated for input data based on a tagging system. Data tags include at least one or more of the following: project tags, requirement tags, task tags, defect tags, personnel tags, and component tags. Establish a relationship between data labels and input data, and write the relationship into the label index table; The system uses a tag index table to retrieve and aggregate input data, forming tagged data.

[0024] Optionally, the above-mentioned versioning of input data based on a preset data model includes: Define a version identifier and version metadata for the input data. The version metadata shall include at least one or more of the following: generation time, data source, change type, and changer identifier. When a change in input data is detected, incremental version data is generated based on the version identifier, while historical version data is retained; Version identifiers, version metadata, and corresponding version data are written to the version repository to enable version tracing and rollback of input data.

[0025] S130. Perform feature processing based on the R&D dataset to obtain feature data.

[0026] In this step, feature extraction, feature encoding, and feature selection can be performed on the R&D dataset in a feature engineering platform to generate feature data. The feature data includes at least one or more of the following: working time features, progress features, blocking event features, holiday impact features, collaborative behavior features, and code quality features.

[0027] Optionally, the above-mentioned working time characteristics include: obtaining working time reporting data, code submission records and / or task status flow records based on task dimension and / or personnel dimension, and calculating actual working time, remaining working time, working time deviation and working time burnout slope, wherein the working time deviation is the difference between actual working time and planned working time, and the working time burnout slope is the rate of change of remaining working time within a preset statistical window.

[0028] Optionally, the above-mentioned progress characteristics include: obtaining the task completion amount and planned completion amount based on the task status flow record and milestone plan, and calculating the progress completion rate, progress deviation, critical path duration and milestone achievement rate, wherein the progress completion rate is the ratio of the task completion amount to the total task, and the progress deviation is the difference between the task completion amount and the planned completion amount.

[0029] Optionally, the above-mentioned blocking event characteristics include: obtaining the type, start time, end time and scope of impact of the blocking event based on the blocking event record, and calculating the blocking duration, blocking frequency, blocking impact and blocking resolution time. The blocking duration is the time difference between the end time and the start time, and the blocking impact is calculated by weighting the number of affected tasks, the priority of affected tasks and the affected working hours.

[0030] Optionally, the above-mentioned holiday impact characteristics include: obtaining project calendar, holiday calendar and personnel scheduling data, calculating the effective number of working days, holiday overlap ratio, available working hours and working day density, wherein the effective number of working days is the number of working days after removing holidays and non-scheduled days in the preset statistical window, and the holiday overlap ratio is the ratio of the number of holiday days to the number of calendar days in the preset statistical window.

[0031] Optionally, the aforementioned collaborative behavior characteristics include: obtaining collaborative event data based on code review records, issue tracking system interaction records, and instant messaging collaboration records, and calculating collaboration frequency, response latency, cross-team collaboration ratio, and rework rate. Here, response latency is the time difference between the time of initiation of a collaborative event and the time of the first response, and cross-team collaboration ratio is the ratio of the number of collaborative events involving at least two teams to the total number of collaborative events.

[0032] Optionally, the above code quality characteristics include: obtaining quality metric data based on static code analysis results, unit test results and defect records output by the continuous integration pipeline, and calculating defect density, number of code smells, test coverage and build failure rate. Among them, defect density is the ratio of the number of defects to the code size within a preset statistical window, and the code size is obtained by measuring the number of lines of code (LOC) and / or the function point size.

[0033] S140. Input the feature data into the artificial intelligence decision engine to execute at least one intelligent processing procedure.

[0034] In this step, feature data may be input into an artificial intelligence decision engine to execute at least one intelligent processing flow, wherein the intelligent processing flow includes at least one of the following: task structure generation, task resource scheduling, progress prediction, risk identification and early warning, document-assisted generation, knowledge accumulation and recommendation.

[0035] In the first possible implementation, the aforementioned intelligent processing flow includes task structure generation. Correspondingly, Figure 2 This is a flowchart illustrating the implementation of S140 according to the first example embodiment of this application. Figure 2 As shown, the task structure generation in S140 includes: S1411. Based on a natural language processing semantic understanding model, perform semantic parsing on the requirement text to extract task entity information.

[0036] Specifically, the process might involve performing sentence segmentation, word segmentation, part-of-speech tagging, and named entity recognition on the requirement text. Based on intent recognition and slot filling, the requirement text would be structurally extracted to obtain task entity information. This task entity information would include at least functional points, priorities, and dependencies. Then, the extracted task entity information would undergo consistency verification, and task entity information that does not meet preset verification rules would be marked as pending confirmation.

[0037] Next, based on a pre-defined function point dictionary and domain terminology database, function point descriptions are identified and function point identifiers are generated. Priority expressions in the requirement text are mapped to pre-defined priority enumeration values ​​based on a priority rule table. Pre-tasks, post-tasks, and parallel relationships in the requirement text are identified based on dependency extraction rules, and dependency edge data is generated to form dependency relationships.

[0038] S1412. Based on the historical similar project template library, perform analogy reasoning on the task entity information to generate a task structure tree.

[0039] Specifically, this can involve acquiring historical project task structure data and feature vectors from a historical similar project template library. A target project feature vector is generated based on the task entity information, and the similarity between the target project feature vector and the historical project feature vectors is calculated. Target templates that meet a preset similarity threshold are selected based on similarity ranking, and task mapping and parameter transfer are performed on the target templates to output analogy reasoning results.

[0040] The calculation of similarity includes: vectorizing the feature vectors of the target project and the feature vectors of historical projects; calculating the similarity score using cosine similarity and / or Euclidean distance; and introducing priority weight coefficients and dependency matching degree weight coefficients into the similarity score to obtain a weighted similarity.

[0041] Furthermore, based on the results of analogical reasoning, the task hierarchy and granularity are determined, and functional points are written into the task structure tree as leaf nodes and / or intermediate nodes. Corresponding nodes are connected as directed edges based on dependencies to form a task structure tree containing hierarchical and dependency relationships. Circular dependency detection and isolated node detection are performed on the task structure tree, and structural correction suggestions are output when abnormal structures are detected.

[0042] S1413. Output the task structure tree as the initial draft of the work breakdown structure.

[0043] Specifically, the task structure tree is converted into a work breakdown structure (WBS) data object. The WBS data object includes at least task identifiers, parent-child relationships, dependency edges, estimated working hours, and a set of candidate responsible persons. The WBS data object is displayed in a tree view at the user interaction layer, generating an editable initial draft of the WBS. Finally, the initial draft of the WBS is associated with and stored along with the corresponding data used for generation, for traceability purposes.

[0044] S1414. Receive user confirmation and / or adjustment operations for the initial draft of the work breakdown structure.

[0045] Specifically, the user interaction layer receives one or more of the following operations for task nodes: adding, deleting, merging, splitting, and moving them hierarchically. Then, it receives editing operations for task attributes, which include at least one or more of the following: task name, priority, dependencies, estimated working hours, and responsible person. Finally, it generates change records for confirmation and / or adjustment operations, and generates an operator identifier and timestamp for each change record.

[0046] S1415. Update the initial draft of the work breakdown structure based on the confirmation operation and / or adjustment operation, and write the updated work breakdown structure back to the data platform.

[0047] Specifically, the confirmation and / or adjustment operations are converted into a structured change instruction sequence. Then, the initial draft of the work breakdown structure is incrementally updated based on the structured change instruction sequence to generate the updated work breakdown structure. The updated work breakdown structure is then subjected to dependency consistency checks and task granularity constraint checks, and conflict warning messages are output when the checks fail.

[0048] Next, the updated work breakdown structure is standardized according to a preset data model, and a structure version identifier is generated. The updated work breakdown structure, structure version identifier, and change records are then written to the data middle platform and synchronized to the data lake warehouse integrated platform. Finally, historical versions are retained based on the structure version identifier to enable version tracking and rollback of the work breakdown structure.

[0049] In the second possible implementation, the aforementioned intelligent processing flow includes task resource scheduling. Correspondingly, Figure 3 This is a flowchart illustrating the implementation of S140 according to the second example embodiment of this application. Figure 3 As shown, task resource scheduling in S140 includes: S1421. Construct a capability profile model for R&D members to obtain capability profile data.

[0050] Specifically, this could involve obtaining one or more of the following: historical task records, code submission records, code review records, defect repair records, and skill certification records of R&D members. These records are then standardized and tagged to generate a set of capability tags. A capability profile model is then trained based on this tag set, and capability profile data is output from the model.

[0051] Optionally, the above capability profile data includes: Skill tags are determined based on historical task records; Calculate code quality score based on code commit history and defect fix history; Response speed is calculated based on code review records; Collaboration scores are calculated based on code review records and collaboration interaction records.

[0052] S1422. Obtain the current load data of the R&D members.

[0053] Specifically, this could involve obtaining the in-transit task list, estimated task hours, task deadlines, and available calendar hours for each development member. Then, based on the in-transit task list, the currently allocated work hours are summarized, and the load factor is calculated based on the available calendar hours. The load factor is the ratio of currently allocated work hours to available calendar hours. Finally, the load factor is associated with and stored in relation to the development member's time window identifier to form the current load data.

[0054] S1423. Based on a multi-objective optimization algorithm, the task and R&D members are matched and solved to comprehensively consider skill matching degree, workload and collaboration history, and output a task allocation scheme.

[0055] Specifically, this can involve constructing task requirement vectors and R&D member capability vectors, and calculating skill matching degree based on these vectors. Candidate R&D members are then filtered for load constraints based on current workload data to obtain a candidate set. Historical collaboration data among R&D members is then acquired, and collaboration affinity is calculated based on this data. Next, a multi-objective optimization algorithm, such as the NSGA-II algorithm, is employed with the objective function of maximizing skill matching degree and collaboration affinity while minimizing workload deviation, to perform the matching solution and output a task allocation scheme.

[0056] Optionally, the above-mentioned skill matching degree calculation includes: The task requirement vector and the R&D member capability vector are vectorized; The matching score is calculated using cosine similarity and / or weighted overlap coefficient. Weighting coefficients are applied to the matching scores based on task priority to obtain the skill matching degree.

[0057] Optionally, the above load constraint filtering includes: Set feasibility constraints for load thresholds and deadlines; When the load rate of a candidate R&D member is greater than the load threshold and / or the available working hours of a candidate R&D member before the task deadline are less than the estimated working hours of the task, the candidate R&D member will be removed from the candidate set.

[0058] Optionally, the calculation of cooperative affinity as described above includes: A collaboration intensity metric is constructed based on the number of times historical tasks were jointly participated in, the number of code review interactions, and the number of cross-team collaborations. Construct collaboration risk indicators based on historical collaboration rework rates and historical collaboration delay rates; The collaboration strength index and the collaboration risk index are weighted and fused together to obtain the collaboration affinity.

[0059] Finally, a sub-configuration reliability score and conflict flag are generated for the task allocation scheme. The conflict flag indicates whether there are constraints such as skill non-compliance, load exceeding threshold constraints, or dependency infeasibility constraints. The task allocation scheme is then displayed at the user interaction layer, and user confirmation and / or adjustment actions are received. Finally, the confirmed task allocation scheme, sub-configuration reliability score, and conflict flag are written back to the data platform.

[0060] S1424. Monitor the status of the task and the status of the R&D team members.

[0061] Specifically, based on the data acquisition bus, task status flow data, remaining task hours and task deadlines are obtained from the project management system; based on the data acquisition bus, code submission frequency, build status and test pass rate are obtained from the code hosting platform and continuous integration platform; based on the data acquisition bus, time entry data, leave data and available calendar hours are obtained from the time system and calendar system.

[0062] Then, the task status flow data, remaining task hours, task deadline, code submission frequency, build status, test pass rate, work hour reporting data, leave data, and available calendar hours are standardized and formed into a monitoring indicator sequence.

[0063] S1425. When a preset rescheduling trigger condition is detected, the matching relationship between tasks and R&D members is re-solved based on a multi-objective optimization algorithm to output an updated task allocation scheme.

[0064] When a preset rescheduling trigger condition is detected, the matching relationship between tasks and R&D members is re-solved based on a multi-objective optimization algorithm to output an updated task allocation scheme. The preset rescheduling trigger condition includes one or more of the following: task delay, R&D member resignation, and R&D member load exceeding a threshold.

[0065] Specifically, the probability of task delay and the number of days of delay are calculated based on the monitoring indicator sequence. The probability of task delay is calculated based on the planned task duration, remaining task hours, and historical completion efficiency. The departure of R&D members is detected based on personnel change event streams. The workload rate of R&D members is calculated based on available calendar hours and allocated hours, and then compared with a preset workload threshold.

[0066] When the probability of a task being delayed is greater than a preset probability threshold and / or the number of days the task is delayed is greater than a preset number of days threshold, the task is deemed to be delayed; when a development member leaves the company, the development member leaves the company; when the load rate is greater than a preset load threshold, the development member load exceeds the threshold, thus triggering a preset rescheduling trigger condition.

[0067] To re-solve the matching relationship between tasks and R&D members based on a multi-objective optimization algorithm, the specific steps are as follows: First, determine the rescheduling scope, which includes the set of target tasks and the set of candidate R&D members affected by preset rescheduling trigger conditions. Then, extract task requirement vectors from the target task set and capability profile data and current load data from the candidate R&D member set. Next, construct a rescheduling constraint set, which includes at least one or more of the following: task dependency constraints, deadline constraints, load threshold constraints, and skill fulfillment constraints. Under the constraints of the rescheduling constraint set, use a multi-objective optimization algorithm to match the target task set and the candidate R&D member set to obtain an updated matching relationship.

[0068] Furthermore, the aforementioned set of rescheduling constraints also includes stability constraints, which are used to limit the number of task changes and / or the number of R&D member changes. The stability constraints are implemented by setting upper limit thresholds for the number of changed tasks and / or the number of changed R&D members.

[0069] Finally, based on the updated matching relationships, a task-member allocation list, adjustment reason identifiers, and impact assessment results are generated. The impact assessment results include at least one or more of the following: changes in expected completion time, changes in load factor, and changes in critical path. The updated task allocation scheme is displayed at the user interaction layer, and user confirmation and / or adjustment operations are received. The confirmed updated task allocation scheme, along with the adjustment reason identifiers and impact assessment results, is written back to the data platform, generating corresponding version identifiers and change records.

[0070] In the third possible implementation, the aforementioned intelligent processing flow includes task resource scheduling. Correspondingly, Figure 4 This is a flowchart illustrating the implementation of S140 according to the third example embodiment of this application. Figure 4 As shown, the progress prediction in S140 includes: S1431. Obtain historical working hours data, current progress data, blocking event data, and holiday impact factors.

[0071] Specifically, historical work time data and real-time work time reporting data are obtained from the work time system based on the data acquisition bus, and the historical work time data is aggregated according to task dimensions and preset statistical granularity.

[0072] The current progress data is obtained from the project management system based on the data acquisition bus. The current progress data includes at least the task status, completion percentage, remaining working hours, and milestone achievement status.

[0073] Blocking event data is acquired from the defect management system and collaboration system based on the data acquisition bus. The blocking event data includes at least the blocking type, blocking start time, blocking end time, and blocking impact range.

[0074] The impact factors of holidays are obtained based on the calendar system. The impact factors of holidays include at least the statutory holiday identifier, weekend identifier, time off in lieu identifier, and available working hours in the team's work calendar.

[0075] Historical working hours data, current progress data, congestion event data, and holiday impact factors are time-aligned and standardized to form a training dataset and prediction input sequence for time series prediction models.

[0076] S1432. Generate completion time prediction results based on time series prediction model.

[0077] Optionally, the time series prediction models mentioned above include one or more of the Prophet model and the Long Short-Term Memory (LSTM) network model.

[0078] Specifically, the target sequence for prediction is determined, which is the sequence of remaining working hours for the task and / or the sequence of remaining workload for the project.

[0079] Input the predicted input sequence into the time series forecasting model, and output the predicted values ​​of remaining working hours and / or remaining workload for multiple future time steps.

[0080] Based on the task duration calendar, the remaining work hours and / or remaining workload forecasts are converted into estimated completion dates to generate completion time forecasts. The task duration calendar is determined based on the impact of holidays.

[0081] S1433. Based on the Bayesian update mechanism, the completion time prediction results are dynamically corrected by combining real-time progress data.

[0082] Specifically, the predicted distribution corresponding to the completion time prediction results is determined as the prior distribution. Real-time progress data and real-time work hour reporting data are acquired within a preset update cycle, and an observation likelihood function is constructed based on the real-time progress data and real-time work hour reporting data.

[0083] The posterior distribution is calculated based on the prior distribution and the observed likelihood function, and the dynamically corrected completion time prediction is obtained by updating the posterior distribution. Finally, the distribution parameters of the posterior distribution are written into the data platform for rolling correction in subsequent update cycles.

[0084] S1434. Output the progress prediction results with confidence intervals.

[0085] Specifically, the completion time confidence interval is calculated based on the posterior distribution at a pre-set confidence level. The completion time confidence interval includes the lower bound completion date and the upper bound completion date.

[0086] Then, the progress prediction results display data is generated, which includes at least the completion time prediction result, the completion time confidence interval, and the update time identifier.

[0087] Next, the progress prediction results are displayed in the user interaction layer, and the progress prediction results are written back to the data platform.

[0088] In the fourth possible implementation, the aforementioned intelligent processing flow includes risk identification and early warning. Correspondingly, Figure 5 This is a flowchart illustrating the implementation of S140 according to the fourth example embodiment of this application. Figure 5 As shown, the risk identification and warning in S140 includes: S1441. Construct a risk knowledge graph.

[0089] Optionally, the risk knowledge graph should at least link personnel nodes, task nodes, component nodes, and external dependency nodes.

[0090] Specifically, it can be based on a data acquisition bus to acquire personnel data, task data, component data, and external dependency data. Among them, personnel data includes at least personnel identification, role, skill tags, and available working hours; task data includes at least task identification, task status, start time, end time, remaining working hours, and dependency relationships; component data includes at least component identification, code repository path, module boundary, and defect records; and external dependency data includes at least supplier identification, dependent deliverable identification, planned delivery time, and actual delivery time.

[0091] Entity alignment and relationship extraction are performed on personnel data, task data, component data, and external dependency data to form a node set that includes at least personnel nodes, task nodes, component nodes, and external dependency nodes, and an edge set that includes at least one or more relationship edges such as responsible, dependent, influential, and associated.

[0092] Next, the node set and edge set are written into the graph database, and timestamps and version identifiers are generated for the node set and edge set to obtain the risk knowledge graph.

[0093] S1442. Anomaly propagation detection is performed on the risk knowledge graph based on graph neural networks to identify potential risk events.

[0094] Specifically, the training samples for the graph neural network are constructed based on the risk knowledge graph. The training samples include a node feature matrix, an adjacency matrix, and risk labels. The node feature matrix includes at least one or more of the following: work time features, progress features, blocking event features, collaborative behavior features, and code quality features. The training samples are trained using one or more of the following models: Graph Convolutional Network (GCN), Graph Attention Network (GAT), and Temporal Graph Neural Network (TPN) to obtain the anomaly propagation detection model.

[0095] The real-time updated risk knowledge graph is input into the anomaly propagation detection model, which outputs node risk scores, edge propagation probabilities, and risk propagation paths to complete the anomaly propagation detection.

[0096] S1443. Generate early warning information based on potential risk events, and push out early warning information and corresponding recommended measures.

[0097] Specifically, potential risk events are identified based on node risk scores and preset risk thresholds, and the risk type, scope of impact, and severity level of these events are determined. The scope of impact includes at least one or more of the following: the set of affected tasks, the set of affected components, and the set of affected personnel. Optionally, the aforementioned potential risk events include one or more of the following: delay risk, resource conflict risk, and technical debt risk.

[0098] Early warning information is generated based on the risk propagation path. The early warning information includes at least the risk type identifier, trigger indicators, risk evidence, predicted impact, and response time limit.

[0099] Based on a pre-defined rule base for suggested measures and / or a historical case base, suggested measures are matched to the early warning information. The suggested measures include at least one or more of the following: resource adjustment suggestions, dependency change suggestions, quality governance suggestions, and milestone adjustment suggestions.

[0100] The system pushes early warning information and suggested measures to target users through the user interaction layer, and writes the early warning information, suggested measures and push records back to the data platform.

[0101] In the fifth possible implementation, the aforementioned intelligent processing flow includes document-assisted generation. Correspondingly, Figure 6 This is a flowchart illustrating the implementation of S140 according to the fifth exemplary embodiment of this application. Figure 6 As shown, document-aided generation in S140 includes: S1451. Obtain contextual data related to the research and development task.

[0102] Optionally, the context data may include at least one or more of the following: interface definition, requirement description, and change log.

[0103] Specifically, it can be based on the data acquisition bus to obtain requirement descriptions from the requirement management system, interface definitions from the interface management system, and change records from the configuration management system, and associate the requirement descriptions, interface definitions, and change records based on task identifiers.

[0104] The requirements description, interface definition, and change record are deduplicated, standardized, and subject to access control. The processed context data is then written to the data platform to form the input context for the large language model.

[0105] S1452. Input the context data into the large language model to generate a technical document draft.

[0106] Specifically, the document structure can be determined based on a preset document template. The document structure should include at least one or more of the following: background description, interface description, data structure, business process, exception handling, and impact of changes.

[0107] The context data is constructed into prompt words according to the document structure, and the prompt words are input into a large language model to output a technical document draft.

[0108] Add source citation markers to the technical document drafts. The source citation markers are used to indicate the source location of the requirement descriptions, interface definitions and / or change records corresponding to each section of the technical document draft.

[0109] S1453. Perform syntax correction and terminology consistency checks on the technical document draft to output the target document.

[0110] Specifically, this could involve performing syntax correction on a draft technical document to obtain a corrected version.

[0111] The terminology database is used to perform terminology consistency checks on the corrected technical documents. The database includes at least one or more of the following: interface field terms, component naming terms, and abbreviation terms. Replacement suggestions are generated for inconsistent terms.

[0112] Receive user confirmation and / or adjustment requests for replacement suggestions, and generate the target document based on the confirmation and / or adjustment requests.

[0113] Then, the target document and source reference identifier are written back to the data platform, and the corresponding version identifier and change record are generated.

[0114] In the sixth possible implementation, the aforementioned intelligent processing flow includes knowledge accumulation and recommendation. Correspondingly, Figure 7 This is a flowchart illustrating the implementation of S140 according to the sixth example embodiment of this application. Figure 7 As shown, the knowledge accumulation and recommendation in S140 includes: S1461. Extract the decision records, problem records, and solution records from the project process in a structured manner and write them into the enterprise-level R&D knowledge graph.

[0115] Specifically, it can acquire decision records, problem records, and solution records from project management systems, defect management systems, code hosting systems, and document systems based on a data acquisition bus, and then associate and aggregate these records based on project identifiers, task identifiers, and timestamps.

[0116] Entity identification and relation extraction are performed on decision records, problem records, and solution records to generate knowledge entries. Each knowledge entry includes at least an entity set and a relation set. The entity set includes at least one or more of the following: personnel entities, task entities, component entities, defect entities, and solution entities. The relation set includes at least one or more of the following: proposal, influence, dependency, resolution, and association.

[0117] Then, the knowledge entries are written into the graph database to form an enterprise-level R&D knowledge graph, and version identifiers and change records are generated for the knowledge entries.

[0118] S1462. Based on collaborative filtering and content matching algorithms, generate knowledge recommendation results from enterprise-level R&D knowledge graphs.

[0119] Specifically, it can involve obtaining the target user identifier, target task identifier, and target component identifier, and then obtaining user behavior data from the data platform based on the target user identifier. The user behavior data includes at least one or more of the following: browsing history, collection history, citation history, and search history.

[0120] The candidate knowledge item set is calculated using collaborative filtering algorithms based on user behavior data. The collaborative filtering algorithms include user-based collaborative filtering algorithms and / or item-based collaborative filtering algorithms.

[0121] Context subgraphs are extracted from the enterprise-level R&D knowledge graph based on target task identifiers and target component identifiers. Then, based on the text vectorization model, the node text of the context subgraphs and the entry text of the candidate knowledge entry set are vector-encoded to calculate the content similarity using a content matching algorithm.

[0122] The collaborative filtering score and content similarity of the candidate knowledge item set are weighted, fused, and ranked to output the knowledge recommendation result. The knowledge recommendation result includes at least the recommended item identifier, the reason for recommendation, the confidence level, and the source path.

[0123] S1463. Push knowledge recommendation results to target users.

[0124] Specifically, the push strategy is determined based on the target user identifier. The push strategy includes at least the push channel, push frequency and push triggering conditions. The push channel includes one or more of the following: in-site messages, email and instant messaging in the user interaction layer.

[0125] When the push trigger conditions are met, knowledge recommendation results are pushed to the target user through the push channel, and a push record is generated. The push record includes at least the push time, push channel, recommendation item identifier, and user interaction results.

[0126] Push notifications and user interaction results are written back to the data platform, and user interaction results are used as feedback data for collaborative filtering and content matching algorithms to dynamically optimize subsequent knowledge recommendation results.

[0127] S150: Output the management results corresponding to the intelligent processing flow and write the management results back to the data platform.

[0128] In this step, the management results corresponding to the intelligent processing flow are output and written back to the data platform for dynamic optimization of subsequent R&D data management.

[0129] Specifically, management result data is generated for the output of the intelligent processing workflow. This management result data includes at least a result identifier, object identifier, result type, result content, confidence level, generation time identifier, and applicable scope identifier. The management result data is then standardized and encapsulated based on a pre-defined result data model to obtain the management results.

[0130] Then, based on the object identifier, the management results are associated with one or more of the following data in the R&D dataset: requirement data, task data, personnel data, and risk data. Access control and audit log generation are performed on the management results, and the results are written to the data platform's result database. Version identifiers and change records are generated for the management results to enable version tracking and rollback.

[0131] Furthermore, it can also obtain user confirmation and / or adjustment operations on management results and generate feedback data, which includes at least the feedback type, feedback content, feedback user identifier, and feedback time identifier.

[0132] Feedback data is written back to the data platform and used as a supervisory signal to update the parameters of the artificial intelligence decision engine and / or the feature selection strategy of the feature engineering platform.

[0133] Furthermore, based on the updated AI decision engine and / or the updated feature selection strategy, intelligent processing is performed on subsequent input data to generate updated management results.

[0134] In this embodiment, by acquiring input data from the R&D management object, the input data is connected to the data platform via a data acquisition bus and stored and managed in the integrated data lake warehouse platform to obtain an R&D dataset. Then, feature processing is performed on the R&D dataset to obtain feature data. The feature data is then input into the artificial intelligence decision engine to execute at least one intelligent processing flow, outputting management results corresponding to the intelligent processing flow, and writing the management results back to the data platform. This enables centralized accumulation and reusability of management results, providing a sustainable and dynamic optimization basis for subsequent R&D data management, and thus enabling dynamic optimization of subsequent R&D data management.

[0135] It is worth noting that the above embodiments may employ a data acquisition bus to uniformly access data from multiple systems, and complete access control, semantic mapping, and master data alignment in the data middle platform. In the integrated data lake warehouse platform, data cleaning, deduplication, standardization, tagging, and versioning are used to construct R&D datasets suitable for analysis and modeling. Then, in the feature engineering platform, feature extraction, feature encoding, and feature selection are performed on the R&D dataset to generate feature data, aligning structured and unstructured R&D data in the feature space. Next, the feature data is input into an artificial intelligence decision engine, which utilizes natural language processing semantic understanding models, time series prediction models, graph neural networks, and multi-objective optimization algorithms to complete task structure generation, progress prediction, risk identification, and task resource scheduling, respectively. The output management results are then standardized and encapsulated according to a preset result data model and written back to the data middle platform to form a result library. This reduces the cost of cross-system data integration, improves the consistency and interpretability of task breakdown and allocation, enhances the accuracy and timeliness of progress forecasting, increases the lead time and coverage of risk identification, shortens the technical document generation cycle and improves terminology consistency, and improves knowledge reuse efficiency, thereby solving the problems of data fragmentation leading to the inability to model and management relying on experience leading to the inability to iterate and optimize.

[0136] Figure 8 This is a schematic diagram illustrating the structure of an artificial intelligence-based R&D data management device according to an example embodiment of this application. Figure 8 As shown, the AI-based R&D data management device 300 provided in this embodiment includes: The acquisition module 310 is used to acquire input data of the R&D management object, wherein the input data includes at least R&D data; Processing module 320 is used to connect the input data to the data platform through the data acquisition bus, and store and manage it in the integrated data lake warehouse platform to obtain the R&D dataset; The processing module 320 is also used to perform feature processing based on the R&D dataset to obtain feature data; The execution module 330 is used to input the feature data into the artificial intelligence decision engine to execute at least one intelligent processing flow; The output module 340 is used to output the management results corresponding to the intelligent processing flow and write the management results back to the data platform.

[0137] Figure 9 This is a schematic diagram of the structure of an electronic device according to an example embodiment of this application. For example... Figure 9 As shown, the electronic device 400 provided in this embodiment includes: a processor 401 and a memory 402; wherein: Memory 402 is used to store computer programs, and the memory may also be flash memory.

[0138] Processor 401 is used to execute the execution instructions stored in the memory to implement the various steps in the above method. For details, please refer to the relevant descriptions in the preceding method embodiments.

[0139] Alternatively, the memory 402 can be either standalone or integrated with the processor 401.

[0140] When the memory 402 is a device independent of the processor 401, the electronic device 400 may further include: Bus 403 is used to connect the memory 402 and the processor 401.

[0141] This embodiment also provides a readable storage medium storing a computer program, which, when executed by at least one processor of an electronic device, enables the electronic device to perform the methods provided in the various embodiments described above.

[0142] This embodiment also provides a program product including a computer program stored in a readable storage medium. At least one processor of an electronic device can read the computer program from the readable storage medium, and the at least one processor executes the computer program to cause the electronic device to perform the methods provided in the various embodiments described above.

[0143] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the claims.

[0144] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A research and development data management method based on artificial intelligence, characterized in that, include: Obtain input data from the R&D management object, wherein the input data includes at least R&D data; The input data is connected to the data platform via a data acquisition bus and stored and managed in the integrated data lake warehouse platform to obtain the R&D dataset. Feature data is obtained by performing feature processing on the aforementioned R&D dataset. The feature data is input into an artificial intelligence decision engine to execute at least one intelligent processing procedure; Output the management results corresponding to the intelligent processing flow, and write the management results back to the data platform.

2. The method according to claim 1, characterized in that, The process of obtaining input data for R&D management objects includes: Receive user requests via text and / or voice commands and / or project process data submitted through the user interaction layer; The voice command is processed to convert speech to text to obtain the required text; The requirement text and the project process data are combined to form the R&D data.

3. The method according to claim 1, characterized in that, The step of connecting the input data to the data platform via a data acquisition bus and storing and managing it in the integrated data lake warehouse platform includes: The input data is subjected to data cleaning, deduplication, standardization, and access control processing. The input data is labeled and versioned based on a preset data model to form the R&D dataset.

4. The method according to claim 1, characterized in that, The feature processing based on the R&D dataset to obtain feature data includes: In the feature engineering platform, feature extraction, feature encoding, and feature selection are performed on the R&D dataset to generate the feature data.

5. The method according to claim 1, characterized in that, The intelligent processing flow includes task structure generation, and the task structure generation includes: The semantic understanding model based on natural language processing is used to perform semantic parsing on the requirement text in order to extract task entity information, which includes at least functional points, priorities and dependencies. The task entity information is compared and reasoned based on a historical similar project template library to generate a task structure tree; Output the task structure tree as a draft of the work breakdown structure.

6. The method according to claim 1, characterized in that, The intelligent processing flow includes task resource scheduling, and the task resource scheduling includes: Build a capability profile model for R&D team members to obtain capability profile data; Obtain the current workload data of R&D members; The task allocation scheme is output by matching tasks with R&D members based on a multi-objective optimization algorithm, taking into account skill matching degree, workload and collaboration history.

7. The method according to claim 1, characterized in that, The intelligent processing flow includes progress prediction, and the progress prediction includes: Obtain historical working hours data, current progress data, congestion event data, and holiday impact factors; Generate completion time prediction results based on time series prediction models; Based on the Bayesian update mechanism, the predicted completion time is dynamically corrected by combining real-time progress data. Output the progress prediction results with confidence intervals.

8. The method according to claim 1, characterized in that, The intelligent processing flow includes risk identification and early warning, and the risk identification and early warning includes: Construct a risk knowledge graph, which is associated with at least personnel nodes, task nodes, component nodes, and external dependency nodes; Anomaly propagation detection is performed on the risk knowledge graph based on graph neural networks to identify potential risk events; Early warning information is generated based on the potential risk events, and the early warning information and corresponding suggested measures are pushed out.

9. The method according to claim 1, characterized in that, The intelligent processing flow includes document-assisted generation, and the document-assisted generation includes: Obtain contextual data related to the research and development task, wherein the contextual data includes at least one or more of the following: interface definition, requirement description, and change record; The context data is input into a large language model to generate a technical document draft; The draft technical document is subjected to syntax correction and terminology consistency checks to output the target document.

10. The method according to claim 1, characterized in that, The intelligent processing flow includes knowledge accumulation and recommendation, and the knowledge accumulation and recommendation includes: The decision-making records, problem records, and solution records during the project process are extracted in a structured manner and written into an enterprise-level R&D knowledge graph; Based on collaborative filtering and content matching algorithms, knowledge recommendation results are generated from the enterprise-level R&D knowledge graph. The knowledge recommendation results are pushed to the target users.