Data research and development demand analysis method and device, equipment, storage medium and product
Patent Information
- Application Number
- CN202610674645.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-15
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]本申请提供一种数据研发需求解析方法、装置、设备、存储介质及产品,用以解决现有技术中缺乏统一的数据治理框架和协同机制,导致需求-开发链路割裂、数据复用性差、模型与业务脱节的技术问题
[0020]本申请提供的数据研发需求解析方法、装置、设备、存储介质及产品,通过自然语言处理模型与历史资产知识库的协同作用,解决了传统需求-开发链路割裂的问题。自然语言处理模型将业务需求转化为结构化任务树,避免人工拆解需求的主观性和低效性;历史资产知识库通过匹配历史案例,提供可复用资产推荐,减少跨团队沟通成本。统一任务调度引擎将任务树与推荐资产列表整合,确保开发流程的连贯性。需求响应周期显著缩短,任务拆解的准确性提升,跨团队协作效率提高,同时减少因需求理解偏差导致的返工风险。
Smart Images

Figure CN122817540A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of artificial intelligence and financial technology, and in particular to a method, apparatus, device, storage medium and product for analyzing data research and development needs. Background Technology
[0002] In the context of digital transformation, enterprise data assets are growing exponentially, and data has become a core resource driving business innovation and decision optimization. Enterprises need to achieve end-to-end processing of data, from raw data collection, cleaning, and modeling to visualization and analysis, through data R&D processes in order to unlock data value and build intelligent capabilities.
[0003] Current data development technologies primarily rely on fragmented toolchains, lacking a unified data governance framework and collaboration mechanism among these tools. During the requirements management phase, business personnel must manually break down requirements through meetings or documents. Tasks such as data development, model training, and report creation are executed across different systems, leading to misunderstandings of requirements and disjointed task schedules. In asset management, assets such as data tables, feature engineering scripts, model versions, and report templates are stored in different databases or file systems, lacking cross-type relational views. This makes it difficult for developers to quickly retrieve and reuse existing resources, resulting in widespread duplication of development. In the development and testing phase, test cases for data, models, and reports must be executed independently in their respective systems, lacking end-to-end simulation verification, leading to data gaps in the production environment. Summary of the Invention
[0004] This application provides a data R&D requirements analysis method, apparatus, equipment, storage medium, and product to address the technical problems in the prior art, such as the lack of a unified data governance framework and collaboration mechanism, which leads to a fragmented requirements-development link, poor data reusability, and a disconnect between models and business.
[0005] Firstly, this application provides a method for analyzing data development needs, including:
[0006] Obtain business requirements;
[0007] By using natural language processing, business requirements are analyzed and key elements are extracted.
[0008] The key elements are matched with historical demand cases in the preset historical asset knowledge base to obtain the matching results;
[0009] Based on the matching results, a structured task tree and a list of reusable assets are generated according to the business requirements through the task scheduling engine.
[0010] Secondly, this application provides a data R&D needs analysis device, comprising:
[0011] The acquisition module is used to acquire business requirements;
[0012] The parsing module is used to parse business requirements through natural language processing and extract key elements;
[0013] The matching module is used to match key elements with historical demand cases in the preset historical asset knowledge base to obtain matching results;
[0014] The generation module is used to generate a structured task tree and a list of reusable assets that correspond to business requirements based on the matching results and through the task scheduling engine.
[0015] Thirdly, this application provides an electronic device, including: a processor and a memory communicatively connected to the processor;
[0016] The memory stores instructions that the computer executes;
[0017] The processor executes computer-executable instructions stored in memory to implement any of the methods of the first aspect.
[0018] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the method of any one of the first aspects.
[0019] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the method of any one of the first aspects.
[0020] The data-driven R&D requirements analysis method, apparatus, equipment, storage medium, and products provided in this application solve the problem of the traditional fragmented requirements-development chain through the synergy of a natural language processing model and a historical asset knowledge base. The natural language processing model transforms business requirements into a structured task tree, avoiding the subjectivity and inefficiency of manual requirements decomposition; the historical asset knowledge base provides reusable asset recommendations by matching historical cases, reducing cross-team communication costs. A unified task scheduling engine integrates the task tree with the recommended asset list, ensuring the continuity of the development process. The requirements response cycle is significantly shortened, the accuracy of task decomposition is improved, cross-team collaboration efficiency is increased, and the risk of rework due to misunderstandings of requirements is reduced. Attached Figure Description
[0021] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0022] Figure 1 A flowchart illustrating the data development process for enterprise evaluation;
[0023] Figure 2A flowchart illustrating a data R&D requirements analysis method provided in this application embodiment;
[0024] Figure 3 A schematic diagram of the structure of a data intelligence R&D system provided in an embodiment of this application;
[0025] Figure 4 A flowchart illustrating the operation of a project management module is provided in this embodiment of the application.
[0026] Figure 5 A flowchart illustrating the operation of an asset exploration module provided in this application embodiment;
[0027] Figure 6 A flowchart illustrating the operation of a research and development testing module is provided in this embodiment of the application.
[0028] Figure 7 A flowchart illustrating the operation of a deployment module is provided in an embodiment of this application.
[0029] Figure 8 A flowchart illustrating the operation of an operation and maintenance management module is provided in this embodiment of the application.
[0030] Figure 9 A schematic diagram of a data R&D requirements analysis device provided in this application embodiment;
[0031] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0032] The accompanying drawings have illustrated specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to specific embodiments. Detailed Implementation
[0033] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0034] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of the relevant data all comply with the relevant laws, regulations, and standards of the relevant countries and regions, have taken necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation access points for users to choose to authorize or refuse.
[0035] Furthermore, the technical solution involved in this application, which involves big data analysis of user information (including but not limited to personal biometrics, identity data, consumption data, asset data, electronic terminal operation data, etc.) and the use of artificial intelligence technology for automated decision-making, and makes decisions that have a significant impact on personal rights based on the results of automated decision-making, provides users with corresponding operation entry points for users to choose to agree to or reject the results of automated decision-making; if the user chooses to reject, the process will proceed to the expert decision-making process.
[0036] It should be noted that the data R&D requirement analysis method, apparatus, equipment, storage medium and product provided in this application can be used in the fields of artificial intelligence and fintech, as well as in any other field. The application fields of the data R&D requirement analysis method, apparatus, equipment, storage medium and product in this application are not limited.
[0037] The specific application scenarios for this application are enterprise-level data development scenarios, such as corporate evaluation in the banking industry, customer profiling analysis in the retail industry, and predictive maintenance of equipment in the manufacturing industry. Taking corporate evaluation in the banking industry as an example... Figure 1 A flowchart illustrating the data development process for enterprise evaluation, such as... Figure 1 As shown, after business personnel put forward their requirements, they need to integrate multi-source data such as customer transaction data and external economic indicators, and through data cleaning, feature engineering, model training, scoring model development and report generation, finally output an evaluation report.
[0038] In existing technologies, data engineers, algorithm engineers, and data analysts must complete their tasks independently in different systems, leading to low collaboration efficiency. Data development processes rely on separate toolchains, lacking a unified data governance framework and collaboration mechanism among these tools. Data engineers process data in extraction, transformation, and loading tools, algorithm engineers develop models on modeling platforms, and analysts design visualizations in reporting tools. Each step relies on manual coordination, resulting in long response cycles, redundant data flow, and repetitive asset development. Furthermore, assets such as data tables, feature engineering scripts, model versions, and report templates are managed in a decentralized manner, lacking cross-type relational views. Developers find it difficult to quickly retrieve and reuse existing resources, leading to widespread duplication of development.
[0039] Therefore, the following core technical problems exist in the existing data R&D process: 1. Fragmented requirement-development chain: In the traditional model, business requirements need to be manually broken down into independent tasks such as data cleaning, model training, and report development, relying on verbal communication or document transmission, which easily leads to misunderstandings of requirements, schedule gaps, and project delays. 2. Poor asset reusability: Assets such as data tables, feature engineering scripts, model versions, and report templates are managed in a scattered manner, lacking cross-type relationship views. Developers cannot quickly identify reusable resources, resulting in redundant development and resource waste. 3. Model disconnect from business: AI model development relies on professional teams, and business personnel cannot directly participate in iterations. The model has a long cycle from requirement to deployment, making it difficult to respond to rapidly changing business needs. 4. Fragmented R&D and testing: The R&D and testing stages of data, models, and reports are independent, lacking end-to-end simulation verification, resulting in data gaps in the production environment. 5. Lack of end-to-end collaboration: Existing technologies lack a unified technical foundation and workflow engine, making it impossible to achieve seamless connection between data exploration, processing, modeling, report development, testing, deployment, and maintenance, as well as end-to-end lineage tracing of data assets.
[0040] This application provides a data R&D requirement analysis method, apparatus, equipment, storage medium, and product that uses a natural language processing model to parse natural language requirements input by business personnel and extract key elements. The extracted elements are then matched with historical requirement cases in a historical asset knowledge base to generate matching results. Based on these matching results, a unified task scheduling engine generates a structured task tree and a list of reusable assets, aiming to solve the aforementioned technical problems of existing technologies.
[0041] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0042] Figure 2 This is a flowchart illustrating a data R&D requirements analysis method provided in an embodiment of this application, as shown below. Figure 2 As shown, the method includes:
[0043] S201. Obtain business requirements.
[0044] In one example, business requirements may include structured or semi-structured business requirement forms, which can be obtained through a smart form interface. This application does not restrict the method of obtaining business requirements.
[0045] S202. Through natural language processing, the business requirements are analyzed and key elements are extracted.
[0046] In one example, natural language processing technology is used to extract key elements from business requirements. These key elements may include data tables, processing logic, delivery methods, etc., and are not limited in this application.
[0047] S203. Match the key elements with historical demand cases in the preset historical asset knowledge base to obtain the matching results.
[0048] In one example, the parsed key elements can be matched with historical requirement cases and related assets in a pre-defined historical asset knowledge base. Based on similarity algorithms (such as keyword matching, semantic vector matching, etc., which are not limited in this application), reusable historical requirement cases and data assets can be recommended. The pre-defined historical asset knowledge base includes a database that stores historical requirement cases and their related assets (such as data tables, model versions, report templates), for example, containing "customer evaluation forms" and their associated feature engineering scripts.
[0049] S204. Based on the matching results, the task scheduling engine generates a structured task tree and a list of reusable assets corresponding to the business requirements.
[0050] In one example, the task scheduling engine may include system modules for coordinating task generation and asset recommendation, such as synchronizing a structured task tree and a list of recommended assets to the development platform. The structured task tree may include a hierarchical task structure that breaks down business requirements into executable steps, such as "data cleaning → feature extraction → model training → report generation". The reusable asset recommendation list may include a list of directly reusable historical assets (such as data tables, scripts, and models), for example, recommending a "customer profile table" instead of the "basic information table" described by the requester.
[0051] In one implementation scenario, a business user submits a request to "generate a customer rating form." The system uses natural language processing technology to extract the keyword "customer rating," matches it with similar historical requirement cases (such as "customer evaluation forms") from a pre-set historical asset knowledge base, recommends reusable feature engineering scripts and model assets, and generates a task tree (such as data cleaning). Feature extraction Model training Report generation).
[0052] In another implementation scenario, the aforementioned data-driven R&D requirement analysis method can be implemented through a data-intelligent R&D system. Figure 3 This is a schematic diagram of the structure of a data intelligence R&D system provided in an embodiment of this application, such as... Figure 3 As shown, the data intelligence R&D system includes: project management module 101, asset exploration module 102, R&D testing module 103, release and deployment module 104, and operation and maintenance management module 105.
[0053] Specifically, Figure 4 This application provides a flowchart illustrating the operation of a project management module, as shown in the embodiments below. Figure 4 As shown, the project management module enables multi-dimensional (data, model, report) requirement management and intelligent parsing. It collects requirements through intelligent forms, providing structured / semi-structured form interfaces to gather business requirements, including data processing, modeling, and report creation. It supports natural language processing to extract key elements (data tables, processing logic, delivery format). The parsed key requirement elements are matched against the platform's knowledge base (which stores historical requirement cases and their associated assets and tasks). Based on similarity algorithms (such as keyword and semantic vector matching), it recommends reusable historical requirement cases and related data assets (tables, scripts), model assets, and report templates, and provides reuse suggestions. It displays data / reports / model development progress by type, automatically associating code submissions, test results, and other development assets. It supports real-time task progress tracking and risk warnings (such as automatically highlighting overdue tasks in red).
[0054] Figure 5 This application provides a flowchart illustrating the operation of an asset exploration module, as shown in the embodiments below. Figure 5 As shown, 2. Asset Exploration Module: Enables full-domain data asset discovery and intelligent recommendation. Multi-source metadata graph construction: Metadata from databases, big data platforms, and application programming interface (API) services is collected through connectors. When developers write structured query language (SCL) code in the data development interface, the SCL parsing engine deeply parses the SCL script, and the system automatically recommends related model feature tables and historical report datasets. When algorithm engineers submit new models, the system intelligently matches and adapts data preprocessing pipelines and reusable feature engineering scripts. When analysts design reports, the system automatically associates the API endpoints and data lineage links of published models, constructing a field-level lineage graph that supports four-layer association tracing: table-field-task-API. After users input their requirements, the intelligent recommendation engine uses a hybrid recommendation algorithm: 50% collaborative filtering (historical usage similarity) + 30% content matching (field semantic similarity) + 20% popularity weighting; it returns the Top 5 candidate assets and reuse guidance in real time (e.g., a "customer profile table" can replace the "basic information table" described by the requester).
[0055] Figure 6 This application provides a flowchart illustrating the operation of a research and development testing module, as shown in the embodiments below. Figure 6As shown, the R&D testing module provides an online integrated development environment and visual orchestration tools, integrating a code template library and AI-assisted generation (such as automatic completion of structured query language). It supports drag-and-drop configuration of data processing pipelines, AI model development pipelines, BI report creation pipelines, etc., automatically generating executable code and supporting manual secondary editing. It generates test cases based on a rule engine and performs data quality checks. It is deeply integrated with a distributed version control system, supporting code difference comparison and branch merge conflict detection. The integrated rule engine automatically generates test cases for data quality checks based on the definition, constraints, and business rules of data assets, supporting automated testing of data processing tasks, model prediction results, and report datasets.
[0056] Figure 7 This application provides a flowchart illustrating the operation of a deployment module, as shown in the embodiments of the present application. Figure 7 As shown, the deployment module enables production deployment and version management. It supports containerized image building, allowing for incremental deployments based on branch lines, user groups, and traffic ratios, with phased rollout according to traffic volume and real-time error rate monitoring. It automatically compares core metrics between old and new versions (e.g., rollback if model accuracy fluctuation exceeds 2%), retains historical version snapshots, and allows for rapid rollback in case of anomalies.
[0057] Figure 8 A flowchart illustrating the operation of an operation and maintenance management module is provided in this embodiment of the application, such as... Figure 8 As shown, the operations and maintenance management module provides comprehensive monitoring and alerts. It collects logs and metrics in real time and uses artificial intelligence to predict capacity bottlenecks. It includes pre-set contingency plans (such as service restarts / node replacements) and root cause analysis to pinpoint problems. It dynamically adjusts cloud resource quotas and automatically reclaims idle instances. Log auditing links operation logs to user identities, supporting traceability and evidence collection.
[0058] The data-driven R&D requirement analysis method provided in this embodiment solves the problem of fragmented traditional requirement-development chains through the synergy of a natural language processing model and a historical asset knowledge base. The natural language processing model transforms business requirements into a structured task tree, avoiding the subjectivity and inefficiency of manual requirement breakdown; the historical asset knowledge base provides reusable asset recommendations by matching historical cases, reducing cross-team communication costs. A unified task scheduling engine integrates the task tree with the recommended asset list, ensuring the continuity of the development process. Requirement response cycles are significantly shortened, the accuracy of task breakdown is improved, cross-team collaboration efficiency is increased, and the risk of rework due to misunderstandings of requirements is reduced.
[0059] Optionally, natural language processing can be used to parse business requirements and extract key elements, including: semantic parsing of business requirements using a pre-set context-aware model, and generating a context-aware structured task tree by combining historical case relationships in a pre-set business knowledge base.
[0060] In one example, the context-aware model may include: a natural language processing model based on a business knowledge base that is updated in real time, capable of parsing by combining contextual semantics and historical case relationships, such as identifying the definitional differences of "customer activity" under different business objectives in a banking scenario. The business knowledge base may include: a database storing industry standard definitions, historical requirement cases, and related assets, such as containing different definitions of "customer activity" in activity and churn warnings.
[0061] For example, when a context-aware model performs semantic parsing of a business requirement description, it combines historical case relationships from the business knowledge base (such as the definition of "customer activity" in different business scenarios) to generate a context-aware structured task tree. For instance, when a business user inputs "generate a score based on transaction data," the model combines current industry standards (such as defining "transaction frequency" as "number of transactions per month") to... (5 times) Generate task tree "Data Cleaning" Feature extraction Model training "Generate reports" and ensure that the task logic is consistent with the business objectives.
[0062] By combining a dynamic context-aware model with real-time updates to the business knowledge base, ambiguity issues caused by polysemous terminology and complex business logic in requirements analysis are resolved. The dynamic context-aware model can automatically match the most appropriate definition based on the current business objectives, thereby reducing misunderstandings of requirements, improving task breakdown efficiency, and lowering cross-team communication costs.
[0063] Optionally, based on the matching results, a structured task tree and a reusable asset recommendation list corresponding to the business requirements are generated through a task scheduling engine. This includes: weighting and sorting the matching results, and generating a reusable asset recommendation list through collaborative filtering recommendation, content matching recommendation, and popularity-weighted recommendation.
[0064] In one example, a recommendation algorithm combining collaborative filtering (based on historical usage similarity), content matching (based on field semantic similarity), and popularity weighting (based on asset usage frequency) can be used, for instance, to recommend a "customer profile table" as an alternative asset to the "basic information table" in a banking scenario. Collaborative filtering recommendations can include recommending assets based on historical usage similarity; for example, if a user frequently uses the "customer profile table," similar assets can be recommended to them.
[0065] For example, when weighting the matching results, a list of asset recommendations is generated by combining collaborative filtering recommendations (based on historical usage similarity), content matching recommendations (based on field semantic similarity), and popularity-weighted recommendations (based on asset usage frequency). For instance, when developers write structured query language scripts, the system uses a hybrid recommendation algorithm to recommend "customer profile tables" (collaborative filtering), "transaction feature scripts" (content matching), and "high-popularity model application programming interfaces" (popularity-weighted), ensuring the accuracy and diversity of the recommendation results.
[0066] By employing a hybrid recommendation algorithm to weight and rank the matching results, the problem of poor asset reusability is addressed. This hybrid algorithm, combining collaborative filtering, content matching, and popularity weighting, reduces redundant development and improves asset reusability.
[0067] Optionally, the method also includes: visually orchestrating the structured task tree through a preset development environment to generate executable code, which supports secondary editing.
[0068] In one example, a structured task tree is visually orchestrated using a low-code development environment to generate executable code that supports manual editing. The low-code development environment includes a development platform that provides visual orchestration tools and code template libraries, such as configuring a data processing pipeline via a drag-and-drop interface. Visual orchestration includes configuring task flows through a graphical interface, such as configuring tasks like "data cleaning". Feature extraction Drag and drop the "Model Training" option onto the development platform. After generating a structured task tree, visualize and orchestrate the task tree using a low-code development environment, for example, by assigning "Data Cleaning" to the task tree. Feature extraction Drag and drop the "Model Training" command into the development platform to generate executable code. Developers can then edit the generated code, such as adjusting feature extraction rules or model parameters, to ensure the task tree aligns with business requirements.
[0069] By using a low-code development environment to visually orchestrate structured task trees, the problem of model development relying on specialized teams is solved. This lowers the barrier to entry for model development and shortens the iteration cycle.
[0070] Optionally, the context-aware model specifically includes at least one of the following: a context-aware model based on semantic vector matching; and / or, a context-aware model based on a rule engine.
[0071] In one example, semantic vector matching could include converting text into vector representations and calculating similarity, such as converting "customer rating" and "customer review" into vectors and calculating their similarity. A rule engine could include context parsing based on predefined rules, such as defining rule matching task logic based on "transaction frequency." Context-aware models can parse descriptions of business requirements through semantic vector matching or rule engines. For example, in a banking scenario, a semantic vector matching model calculates the similarity between "customer activity" and historical cases, while a rule engine uses predefined rules (such as "transaction frequency") to perform context parsing. The matching task logic (5 times / month) ensures that the parsing results are consistent with the business objectives.
[0072] Dynamic context awareness is achieved through semantic vector matching or rule engines, enhancing the flexibility of requirement parsing. The semantic vector matching model calculates the similarity to historical cases, while the rule engine, based on predefined rule matching task logic, ensures that the parsing results are consistent with business objectives, thereby reducing misunderstandings of requirements.
[0073] Optionally, the default development environment specifically includes at least one of the following: a development environment integrating an artificial intelligence model factory; and / or a development environment integrating a visualization reporting tool.
[0074] In one example, the AI model factory includes a development environment that provides predefined model templates and an automated training process, such as selecting a logistic regression model for training via a drag-and-drop interface. The visualization reporting tool may include a graphical interface that supports report design, such as generating customer profile reports by dragging and dropping fields. Integrating the AI model factory and the visualization reporting tool through a low-code development environment—for example, allowing business personnel to select predefined model templates (such as random forests) for training via the AI model factory and design customer profile reports via the visualization reporting tool—ensures a seamless development process.
[0075] By integrating an AI model factory and visual reporting tools, the ability of business personnel to directly participate in model development and report design has been enhanced. Through the AI model factory, predefined model templates can be selected for training, and customer profile reports can be designed using visual reporting tools. Business personnel can directly adjust model parameters and view scoring reports in real time, thereby shortening the model iteration cycle and ensuring the continuity of the development process.
[0076] Optionally, the method also includes: managing the version of the executable code through containerized deployment; and, during the release and deployment phase, launching the new version of the executable code based on traffic ratio.
[0077] In one example, the containerized deployment module may include a deployment system that supports building containerized images. Proportional rollout may include releasing new versions gradually based on user groups or traffic proportions, such as directing 50% of traffic to a new version of the model service. After generating executable code, a container image is built using the containerized deployment module, and the new version is rolled out gradually based on traffic proportions during the release and deployment phase. For example, the new version of the model service might initially run with 50% traffic, and if performance is stable, it could be gradually scaled up to 100%, ensuring the stability of version updates.
[0078] By deploying modules in a containerized manner and adopting a phased rollout strategy based on traffic volume, the stability of version updates was ensured. New version model services can be rolled out gradually based on traffic volume, avoiding the risk of business interruption due to model performance fluctuations and ensuring the stability of version updates.
[0079] Optionally, the method also includes: tracing data dependencies across platforms in a structured task tree through multi-source heterogeneous data lineage tracing.
[0080] In one example, multi-source heterogeneous data lineage tracing may include a system module that tracks the lineage relationships of cross-platform data sources, such as recording the flow path of the "customer rating" field between platform A and platform B. After generating a structured task tree, the multi-source heterogeneous data lineage tracing module performs cross-platform tracing of data dependencies in the task tree. For example, the "customer rating" field, which the model training task depends on, may come from the "customer table" in platform A and the "transaction log" in platform B. The system constructs cross-platform links through the lineage tracing module, supporting field-level tracing.
[0081] The multi-source heterogeneous data lineage tracking module enables cross-platform data dependency tracing, resolving the issue of model failure caused by changes in data sources. If the data format is adjusted, the lineage tracking module can promptly identify the change and pinpoint the source of the problem, preventing model failure.
[0082] Figure 9 A schematic diagram of a data R&D requirements analysis device provided in this application embodiment is shown below. Figure 9 As shown, the data R&D requirement analysis device 90 provided in this embodiment includes:
[0083] Module 901 is used to obtain business requirements;
[0084] The parsing module 902 is used to parse business requirements through natural language processing and extract key elements;
[0085] Matching module 903 is used to match key elements with historical demand cases in the preset historical asset knowledge base to obtain matching results;
[0086] The generation module 904 is used to generate a structured task tree and a list of reusable assets corresponding to business requirements based on the matching results through the task scheduling engine.
[0087] In one possible implementation, the parsing module 902 is specifically used to: perform semantic parsing of business requirements through a preset context-aware model, and generate a context-aware structured task tree by combining the historical case relationships in a preset business knowledge base.
[0088] In one possible implementation, the generation module 904 is specifically used to: weight and sort the matching results, and generate a reusable asset recommendation list through collaborative filtering recommendation, content matching recommendation and popularity weighted recommendation.
[0089] In one possible implementation, the data R&D requirement analysis device is also specifically used to: visualize and orchestrate a structured task tree through a preset development environment, generate executable code, and the executable code supports secondary editing.
[0090] In one possible implementation, the context-aware model specifically includes at least one of the following: a context-aware model based on semantic vector matching; and / or, a context-aware model based on a rule engine.
[0091] In one possible implementation, the pre-defined development environment specifically includes at least one of the following: a development environment integrating an artificial intelligence model factory; and / or a development environment integrating a visualization reporting tool.
[0092] In one possible implementation, the data-driven R&D requirements analysis device is also specifically used for: managing the version of executable code through containerized deployment; and, during the release and deployment phase, launching the new version of the executable code based on traffic ratios.
[0093] In one possible implementation, the data R&D requirement analysis device is also specifically used to: trace data dependencies in a structured task tree across platforms by tracing the lineage of multi-source heterogeneous data.
[0094] The data R&D requirement analysis device provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.
[0095] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 10As shown, the electronic device 100 may include a memory 1001 and a processor 1002. Optionally, the electronic device may also include a transceiver 1003, wherein the memory 1001 and the processor 1002 communicate; for example, the memory 1001, the processor 1002 and the transceiver 1003 may communicate via a communication bus 1004, the memory 1001 is used to store a computer program, and the processor 1002 executes the computer program to implement the method of the above embodiments.
[0096] Optionally, the aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps in the method embodiments disclosed in this application can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.
[0097] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the methods in any of the above method embodiments.
[0098] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the methods in any of the above method embodiments.
[0099] All or part of the steps in the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a readable memory. When the program is executed, it performs the steps of the above method embodiments; and the aforementioned memory (storage medium) includes: read-only memory (ROM), RAM, flash memory, hard disk, solid-state drive, magnetic tape, floppy disk, optical disk, and any combination thereof.
[0100] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processing unit of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processing unit of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0101] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0102] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0103] Obviously, those skilled in the art can make various modifications and variations to the embodiments of this application without departing from the spirit and scope of this application. Therefore, if these modifications and variations to the embodiments of this application fall within the scope of the claims of this application and their equivalents, this application also intends to include these modifications and variations.
[0104] In this application, the term "comprising" and its variations can refer to non-limiting inclusion; the term "or" and its variations can refer to "and / or". The terms "first", "second", etc., in this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. In this application, "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0105] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.
[0106] It should be further noted that although the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0107] It should be understood that the above-described device embodiments are merely illustrative, and the device of this application can also be implemented in other ways. For example, the division of units / modules in the above embodiments is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units, modules, or components may be combined, or integrated into another system, or some features may be ignored or not executed.
[0108] Furthermore, unless otherwise specified, the functional units / modules in the various embodiments of this application can be integrated into one unit / module, or each unit / module can exist physically separately, or two or more units / modules can be integrated together. The integrated units / modules described above can be implemented in hardware or as software program modules.
[0109] When integrated units / modules are implemented in hardware, the hardware can be digital circuits, analog circuits, etc. The physical implementation of the hardware structure includes, but is not limited to, transistors, memristors, etc. Unless otherwise specified, the processor can be any suitable hardware processor, such as a CPU, GPU, FPGA, DSP, and ASIC, etc. Unless otherwise specified, the storage unit can be any suitable magnetic or magneto-optical storage medium, such as Resistive Random Access Memory (RRAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), Enhanced Dynamic Random Access Memory (EDRAM), High-Bandwidth Memory (HBM), Hybrid Memory Cube (HMC), etc.
[0110] If the integrated unit / module is implemented as a software program module and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0111] In the above embodiments, the descriptions of each embodiment have their own emphasis. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.
[0112] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.
[0113] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A method for analyzing data development requirements, characterized in that, The method includes: Obtain business requirements; The business requirements are analyzed using natural language processing to extract key elements; The key elements are matched with historical demand cases in a preset historical asset knowledge base to obtain matching results; Based on the matching results, a structured task tree and a list of reusable assets are generated through the task scheduling engine to meet the business requirements.
2. The method according to claim 1, characterized in that, The process involves parsing the business requirements using natural language processing to extract key elements, including: By using a pre-set context-aware model, the business requirements are semantically parsed, and combined with the historical case relationships in the pre-set business knowledge base, a context-aware structured task tree is generated.
3. The method according to claim 1, characterized in that, Based on the matching results, the process generates a structured task tree and a recommended list of reusable assets corresponding to the business requirements through a task scheduling engine, including: The matching results are weighted and sorted, and the reusable asset recommendation list is generated through collaborative filtering recommendation, content matching recommendation, and popularity-weighted recommendation.
4. The method according to claim 1, characterized in that, The method further includes: The structured task tree is visualized and arranged using a pre-defined development environment to generate executable code, which supports secondary editing.
5. The method according to claim 2, characterized in that, The context-aware model specifically includes at least one of the following: Context-aware model based on semantic vector matching; And / or, a context-aware model based on a rules engine.
6. The method according to claim 4, characterized in that, The preset development environment specifically includes at least one of the following: An integrated development environment for an AI model factory; And / or, a development environment that integrates visual reporting tools.
7. The method according to claim 6, characterized in that, The method further includes: The executable code is version-managed through containerized deployment; During the release and deployment phase, the new version corresponding to the executable code will be launched based on the traffic ratio.
8. The method according to any one of claims 2-7, characterized in that, The method further includes: By tracing the lineage of multi-source heterogeneous data, cross-platform tracing of data dependencies in the structured task tree is performed.
9. A data R&D requirement analysis device, characterized in that, The device includes: The acquisition module is used to acquire business requirements; The parsing module is used to parse the business requirements through natural language processing and extract key elements; The matching module is used to match the key elements with historical demand cases in the preset historical asset knowledge base to obtain matching results; The generation module is used to generate a structured task tree and a list of reusable assets corresponding to the business requirements based on the matching results and through the task scheduling engine.
10. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1 to 8.
12. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method of any one of claims 1 to 8.