Data generation method, apparatus, device, and medium

By acquiring data from multiple data sources and performing multi-dimensional scoring and causal analysis, combined with a time recurrent neural network model and semantic feature matching, accurate target data is generated, solving the problems of low data application recognition efficiency and inaccurate matching, and achieving efficient and reliable data processing.

CN122132602APending Publication Date: 2026-06-02INDUSTRIAL AND COMMERCIAL BANK OF CHINA

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INDUSTRIAL AND COMMERCIAL BANK OF CHINA
Filing Date
2025-08-25
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

In existing technologies, data application identification efficiency is low, data and business matching is inaccurate, and data analysis accuracy is insufficient, resulting in significant information bias.

Method used

By acquiring data containing preset index terms from multiple data sources, performing multi-dimensional scoring, causal analysis, and matching degree scoring, target data is generated. Then, a time recurrent neural network model is used for trend prediction and semantic feature matching to achieve precise hierarchical processing of the data.

Benefits of technology

It improves the comprehensiveness and accuracy of data acquisition, reduces the risk of information bias, saves computing resources, and improves data processing efficiency and the reliability of target data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122132602A_ABST
    Figure CN122132602A_ABST
Patent Text Reader

Abstract

This disclosure provides a data generation method. It can be applied to the fields of artificial intelligence, big data, and the application of large models in fintech. The method includes: obtaining *a* primary data points containing preset index terms from *m* data sources; performing multi-dimensional scoring on the *a* primary data points to generate *a* dimensional scores; obtaining *b* secondary data points that exceed a first preset threshold; performing causal analysis on the *b* secondary data points to generate *c* tertiary data points; obtaining historical business data and matching the historical business data with the *c* tertiary data points to generate matching scores for the *c* tertiary data points; performing a weighted calculation of the matching score and dimensional score for each of the *c* tertiary data points to generate a feasibility score for the *c* tertiary data points; and obtaining *d* tertiary data points that exceed a second preset threshold to generate *d* target data points. This disclosure also provides a data generation apparatus, device, storage medium, and program product.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the fields of big data technology and artificial intelligence technology, specifically to the application of big data models in the field of financial technology, and particularly to a data generation method, apparatus, device, medium and program product. Background Technology

[0002] With the rapid development of the data economy, enterprises, including banks, are identifying potential application scenarios for their technologies by acquiring and analyzing cutting-edge innovative technologies. Enterprises need to quickly identify cutting-edge technology data and comprehensively assess its application potential; accurately identify and systematically analyze the application possibilities of cutting-edge technology data in their original and cross-domain fields; and accurately and automatically match cutting-edge technology data with potential application scenarios, and evaluate the feasibility of technology application.

[0003] The technologies currently available suffer from low efficiency in data application and identification, inaccurate matching of data with business needs, and insufficient accuracy in data analysis. Furthermore, they have significant limitations in data acquisition and are prone to information bias. Summary of the Invention

[0004] In view of the above problems, this disclosure provides a data generation method, apparatus, device, medium and program product.

[0005] According to a first aspect of this disclosure, a data generation method is provided, comprising: acquiring a first data points containing preset index terms from m data sources, wherein a is greater than 1 and a is an integer, m is greater than 1 and m is an integer; performing multi-dimensional scoring on the a first data points to generate a first data dimensional scores; acquiring first data points with a first data dimensional scores higher than a first preset threshold to generate b second data points, wherein b is less than or equal to a, b is greater than 0 and b is an integer; performing causal analysis on the b second data points to generate c third data points, wherein c is greater than 0 and c is an integer; acquiring historical business data, matching the historical business data with the c third data points to generate matching degree scores for the c third data points; performing weighted calculation on the matching degree score and dimensional score of each of the c third data points to generate a feasibility score for the c third data points; and acquiring d third data points with feasibility scores higher than a second preset threshold to generate d target data points, wherein d is less than or equal to c, d is greater than 0 and d is an integer.

[0006] According to an embodiment of this disclosure, obtaining a first data points containing preset index terms from m data sources includes: obtaining m retrieval rules from the m data sources; performing rule-based processing on the preset index terms based on the m retrieval rules to generate n retrieval expressions corresponding to the m data sources, where n is greater than or equal to m and n is an integer; obtaining e first data points containing preset index terms from the m data sources based on the n retrieval expressions; and performing data preprocessing on the e first data points to generate a first data points, where e is greater than or equal to a and e is an integer.

[0007] According to embodiments of this disclosure, multi-dimensional scoring is performed on the a first data points to generate a first data point dimension scores, including: extracting text features from the a first data points to generate feature keywords for the a first data points; calculating the term frequency inverse document frequency (TNF) of the a first data points based on the feature keywords of the a first data points; performing trend prediction on the a first data points based on a time recurrent neural network model to generate trend prediction values ​​for the a first data points; and performing a weighted calculation of the TNF and the corresponding trend prediction value for each of the a first data points to generate a first data point dimension scores.

[0008] According to an embodiment of this disclosure, causal analysis is performed on b second data to generate c third data, including: segmenting each of the b second data into sentences to generate causal conjunctions and grammatical logic for the b second data; generating b related data based on the causal conjunctions and grammatical logic of the b second data; and concatenating the b related data to generate c third data.

[0009] According to embodiments of this disclosure, the associated data includes: interrelated technical means, technical problems, and application scenarios. The process of concatenating the b associated data to generate c third data includes: merging the associated data of the b associated data that share the same technical means to generate f first sub-associated data of different technical means, wherein each first sub-associated data includes the same technical means and all corresponding technical problems and application scenarios, where f is greater than or equal to 0, f is less than or equal to b, and f is an integer; merging the associated data of the b associated data that share the same technical problems to generate g second sub-associated data of different technical problems, wherein each first sub-associated data includes the same technical means, and all corresponding technical problems and application scenarios. The technical problem, and all corresponding technical means and application scenarios, wherein g is greater than or equal to 0, g is less than or equal to b and g is an integer; the b associated data are merged to generate h third sub-associated data with different application scenarios, wherein each third sub-associated data includes the same application scenario and all corresponding technical means and technical problems, wherein h is greater than or equal to 0, h is less than or equal to b and h is an integer; and the f first sub-associated data with different technical means, the g second sub-associated data with different technical problems and the h third sub-associated data with different application scenarios are merged and deduplicated to generate c third data, wherein the sum of f, g and h is greater than or equal to c.

[0010] According to an embodiment of this disclosure, acquiring historical business data and matching the historical business data with c third data to generate a matching score for the c third data includes: acquiring historical business data; extracting semantic feature data from the historical business data to generate a first semantic feature; extracting semantic features from the c third data to generate a second semantic feature for the c third data; and performing similarity matching scoring on the first semantic feature and the second semantic feature of the c third data respectively to generate a matching score for the c third data.

[0011] According to an embodiment of this disclosure, the method further includes: generating a feasibility report based on the d target data and sending it.

[0012] According to a second aspect of this disclosure, a data generation apparatus is provided, comprising: a first acquisition module, configured to acquire a first data points containing preset index terms from m data sources, wherein a is greater than 1 and a is an integer, and m is greater than 1 and m is an integer; a first generation module, configured to perform multi-dimensional scoring on the a first data points to generate a first data dimensional scores; a second generation module, configured to acquire first data points with a first data dimensional scores higher than a first preset threshold to generate b second data points, wherein b is less than or equal to a, b is greater than 0 and b is an integer; and a third generation module, configured to perform multi-dimensional scoring on the b second data points. The system performs causal analysis to generate c third-party data, where c is greater than 0 and is an integer; a fourth generation module is used to acquire historical business data and match the historical business data with the c third-party data to generate matching scores for the c third-party data; a fifth generation module is used to perform weighted calculations on the matching scores and dimension scores of each of the c third-party data to generate feasibility scores for the c third-party data; and a sixth generation module is used to acquire d third-party data whose feasibility scores are higher than a second preset threshold, and generate d target data, where d is less than or equal to c, d is greater than 0 and is an integer.

[0013] According to an embodiment of this disclosure, the first acquisition module includes: a second acquisition module, configured to acquire m retrieval rules from m data sources; a seventh generation module, configured to perform rule-based processing on the preset index terms based on the m retrieval rules to generate n retrieval expressions corresponding to the m data sources, wherein n is greater than or equal to m and n is an integer; a third acquisition module, configured to acquire e first data items containing preset index terms from the m data sources based on the n retrieval expressions; and an eighth generation module, configured to perform data preprocessing on the e first data items to generate a first data items, wherein e is greater than or equal to a and e is an integer.

[0014] According to an embodiment of this disclosure, the first generation module includes: a ninth generation module, used to extract text features from a first data set to generate feature keywords for the a first data set; a first calculation module, used to calculate the term frequency inverse document frequency (TNF) of the text of the a first data set based on the feature keywords of the a first data set; a tenth generation module, used to perform trend prediction on the a first data set based on a time recurrent neural network model to generate trend prediction values ​​for the a first data set; and an eleventh generation module, used to perform weighted calculation on the TNF and the corresponding trend prediction value of each first data set in the a first data set to generate a dimensional score for the a first data set.

[0015] According to an embodiment of this disclosure, the third generation module includes: a twelfth generation module, used to perform sentence segmentation processing on each of the b second data to generate causal conjunctions and grammatical logic for the b second data; a thirteenth generation module, used to generate b related data based on the causal conjunctions and grammatical logic of the b second data; and a fourteenth generation module, used to concatenate the b related data to generate c third data.

[0016] According to an embodiment of this disclosure, the fourteenth generation module includes: a fifteenth generation module, configured to merge the associated data with the same technical means from the b associated data to generate f first sub-associated data with different technical means, wherein each first sub-associated data includes the same technical means and all corresponding technical problems and application scenarios, wherein f is greater than or equal to 0, f is less than or equal to b, and f is an integer; and a sixteenth generation module, configured to merge the associated data with the same technical problems from the b associated data to generate g second sub-associated data with different technical problems, wherein each first sub-associated data includes the same technical problem and all corresponding technical means and application scenarios. Wherein, g is greater than or equal to 0, g is less than or equal to b, and g is an integer; the seventeenth generation module is used to merge the associated data of the same application scenario in the b associated data to generate h third sub-associated data of different application scenarios, wherein each third sub-associated data includes the same application scenario and all corresponding technical means and technical problems, wherein h is greater than or equal to 0, h is less than or equal to b, and h is an integer; and the eighteenth generation module is used to merge and deduplicate the first sub-associated data of f different technical means, the second sub-associated data of g different technical problems, and the third sub-associated data of h different application scenarios to generate c third data, wherein the sum of f, g, and h is greater than or equal to c.

[0017] According to an embodiment of this disclosure, the fourth generation module includes: a nineteenth generation module, used to acquire historical business data, extract semantic feature data from the historical business data, and generate a first semantic feature; a twentieth generation module, used to extract semantic features from c third data, and generate a second semantic feature of c third data; and a twenty-first generation module, used to perform similarity matching scoring on the first semantic feature and the second semantic feature of c third data respectively, and generate a matching score for c third data.

[0018] According to a third aspect of this disclosure, an electronic device is provided, comprising: one or more processors; and a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors perform the data generation method described above.

[0019] According to a fourth aspect of this disclosure, a computer-readable storage medium is provided that stores executable instructions or a computer program thereon, which, when executed by a processor, cause the processor to perform the data generation method described above.

[0020] According to a fifth aspect of this disclosure, a computer program product is also provided, including a computer program that, when executed by a processor, implements the above-described data generation method.

[0021] This disclosure improves the comprehensiveness of acquired data and reduces the risk of information bias by acquiring data from multiple sources. Then, through a hierarchical processing mechanism of data a→b→c→d, low-value data is eliminated early, improving overall processing efficiency. Furthermore, a dual-threshold matching mechanism with business data enables precise stratification of data value, making the acquired target data more accurate. Overall, it achieves the technical effect of saving computer resources while improving the accuracy and reliability of target data. It can solve the technical problems of low data application identification efficiency, inaccurate data-business matching, and insufficient data analysis accuracy under the current model. Attached Figure Description

[0022] The foregoing contents, as well as other objects, features, and advantages of this disclosure, will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:

[0023] Figure 1 This diagram illustrates an application scenario of the data generation method and apparatus according to embodiments of the present disclosure.

[0024] Figure 2 A flowchart illustrating a data generation method according to an embodiment of the present disclosure is shown schematically;

[0025] Figure 3 A flowchart illustrating the generation of first data in a data generation method according to an embodiment of the present disclosure is shown schematically.

[0026] Figure 4 This illustration schematically shows a flowchart of generating a first data dimension score in a data generation method according to an embodiment of the present disclosure;

[0027] Figure 5 A flowchart illustrating the generation of third data in a data generation method according to an embodiment of the present disclosure is shown schematically.

[0028] Figure 6A flowchart illustrating the data generation method according to an embodiment of the present disclosure is shown.

[0029] Figure 7 A flowchart illustrating the generation of a matching score for third data in a data generation method according to an embodiment of the present disclosure is shown in the schematic diagram.

[0030] Figure 8 A schematic block diagram of a data generation apparatus according to an embodiment of the present disclosure is shown; and

[0031] Figure 9 A block diagram schematically illustrates an electronic device suitable for implementing a data generation method according to an embodiment of the present disclosure. Detailed Implementation

[0032] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.

[0033] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0034] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0035] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).

[0036] The accompanying drawings show some block diagrams and / or flowcharts. It should be understood that some blocks or combinations thereof in the block diagrams and / or flowcharts can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable control device, so that when executed by the processor, these instructions can create means for implementing the functions / operations described in these block diagrams and / or flowcharts.

[0037] First, the technical terms used in this article are explained as follows:

[0038] LSTM (Long Short-Term Memory) is a type of recurrent neural network used to process and predict long-term dependencies in time series data, and it can solve the gradient vanishing problem of traditional neural networks. Data is input into an LSTM and the output is the SHAP value.

[0039] SHAP values ​​(SHapley Additive exPlanations) are a game theory-based interpretability method used to quantify the contribution of each feature in a machine learning model to the prediction result.

[0040] This disclosure provides a data generation method, comprising: acquiring a first data points containing preset index terms from m data sources, wherein a is greater than 1 and is an integer, and m is greater than 1 and is an integer; performing multi-dimensional scoring on the a first data points to generate a first data dimensional scores; acquiring a first data points with dimensional scores higher than a first preset threshold to generate b second data points, wherein b is less than or equal to a, b is greater than 0 and is an integer; performing causal analysis on the b second data points to generate c third data points, wherein c is greater than 0 and is an integer; acquiring historical business data and matching the historical business data with the c third data points to generate matching degree scores for the c third data points; performing a weighted calculation on the matching degree score and dimensional score of each of the c third data points to generate a feasibility score for the c third data points; and acquiring d third data points with feasibility scores higher than a second preset threshold to generate d target data points, wherein d is less than or equal to c, d is greater than 0 and is an integer.

[0041] According to the embodiments of this disclosure, by acquiring data from multiple sources, the comprehensiveness of the acquired data is improved, and the risk of information bias is reduced. Then, through a hierarchical processing mechanism of data a→data b→data c→data d, low-value data is eliminated early, improving overall processing efficiency. Furthermore, through a dual-threshold matching mechanism with business data, precise stratification of data value is achieved, making the acquired target data more accurate. Overall, this achieves the technical effect of saving computer resources while improving the accuracy and reliability of target data. It can solve the technical problems of low data application identification efficiency, inaccurate data-business matching, and insufficient data analysis accuracy in the current model.

[0042] Figure 1 The diagram illustrates an application scenario of the data generation method and apparatus according to embodiments of the present disclosure. It should be noted that... Figure 1 The examples shown are merely examples of scenarios in which the embodiments of this disclosure can be applied, to help those skilled in the art understand the technical content of this disclosure, but do not mean that the embodiments of this disclosure cannot be used in other devices, systems, environments or scenarios.

[0043] like Figure 1 As shown, application scenario 100 according to this embodiment may include a data generation application scenario. Network 104 serves as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables, etc.

[0044] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).

[0045] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0046] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0047] It should be noted that the data generation method provided in this embodiment can generally be executed by server 105. Correspondingly, the data generation apparatus provided in this embodiment can generally be located in server 105. The data generation method provided in this embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the data generation apparatus provided in this embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105.

[0048] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0049] The following will be based on Figure 1 The described scene, through Figures 2-7 The data generation method of the disclosed embodiments is described in detail. It should be noted that the above application scenarios are shown only to facilitate understanding of the spirit and principles of this disclosure, and the implementation of this disclosure is not limited in any way. Rather, the implementation of this disclosure can be applied to any applicable scenario.

[0050] Figure 2 A flowchart illustrating a data generation method according to an embodiment of the present disclosure is shown schematically.

[0051] like Figure 2 As shown, the method 200 includes steps S201 to S207.

[0052] Step S201: Obtain a first data points containing preset index terms from m data sources, where a is greater than 1 and a is an integer, and m is greater than 1 and m is an integer.

[0053] Figure 3 The flowchart illustrating the generation of first data in a data generation method according to an embodiment of the present disclosure is shown schematically.

[0054] like Figure 3 As shown, the method 300 includes steps S301 to S304.

[0055] Step S301: Obtain m search rules from m data sources.

[0056] For example, the m data sources include major academic databases, preprint platforms, and numerous other academic platforms.

[0057] Step S302: Based on the m retrieval rules, perform rule-based processing on the preset index terms to generate n retrieval expressions corresponding to the m data sources, where n is greater than or equal to m and n is an integer.

[0058] For example, major academic databases, preprint platforms, and numerous academic platforms all have their own search rules, and n search queries are generated based on these search rules.

[0059] Step S303: Based on the n search terms, obtain e first data points containing preset index terms from m data sources.

[0060] For example, preset index terms may include: index terms related to emerging and innovative technologies. Based on n search queries, e first data points containing index terms related to emerging and innovative technologies are retrieved from m data sources.

[0061] Step S304: Perform data preprocessing on the e first data to generate a first data, where e is greater than or equal to a, and e is an integer.

[0062] For example, data preprocessing of the e first data may include: deleting duplicate data, handling missing values, handling outliers, or smoothing noisy data.

[0063] By standardizing multi-source retrieval data, the relevance of the data was improved, the data quality was enhanced, and data preprocessing further improved data processing efficiency and saved CPU resources.

[0064] Return to reference Figure 2 In step S202, the a first data points are scored in multiple dimensions to generate a first data point dimension scores.

[0065] Figure 4 The flowchart illustrating the generation of a first data dimension score in a data generation method according to an embodiment of the present disclosure is shown in the illustration.

[0066] like Figure 4 As shown, the method 400 includes steps S401 to S404.

[0067] Step S401: Extract text features from a first data points to generate feature keywords for the a first data points.

[0068] For example, by performing preprocessing such as word segmentation, part-of-speech tagging, and entity recognition on the documents, text features can be extracted from the first data points (a).

[0069] Step S402: Based on the feature keywords of the a first data, calculate the term frequency inverse document frequency of the a first data text.

[0070] Step S403: Based on the time recurrent neural network model, perform trend prediction on the a first data points to generate trend prediction values ​​for the a first data points.

[0071] For example, the SHAP value of the first data can be output by LSTM as the trend prediction value of the first data.

[0072] Step S404: Perform a weighted calculation on the word frequency inverse document frequency and the corresponding trend prediction value for each of the a first data points to generate a first data dimensional scores.

[0073] For example, based on the application scenario, a first weight and a second weight are preset. The product of the inverse document frequency of the word frequency of each of the a first data points and the first weight is calculated as a first association value. The product of the trend prediction value and the second weight in the a first data points is calculated as a second association value. The sum of the first association value of the inverse document frequency of the word frequency of each of the a first data points and the second association value of the corresponding trend prediction value is calculated to generate a score for each dimension of the a first data points.

[0074] By generating a first data dimension score using word frequency inverse document frequency and the corresponding trend prediction value, the accuracy and reliability of the first data dimension score can be improved.

[0075] Return to reference Figure 2 In step S203, first data with scores of a first data dimension that are higher than the first preset threshold are obtained, and b second data are generated, wherein b is less than or equal to a, b is greater than 0 and b is an integer.

[0076] Step S204: Perform causal analysis on b second data points to generate c third data points, where c is greater than 0 and c is an integer.

[0077] Figure 5 A flowchart illustrating the generation of third data in a data generation method according to an embodiment of the present disclosure is shown.

[0078] like Figure 5 As shown, the method 500 includes steps S501 to S503.

[0079] Step S501: Perform sentence segmentation on each of the b second data to generate causal conjunctions and grammatical logic for the b second data.

[0080] For example, by using a semantic big data model or natural language processing algorithm to segment each of the b second data texts to be processed, causal conjunctions and grammatical logic for the b second data can be generated.

[0081] Step S502: Based on the causal conjunctions and grammatical logic of the b second data, generate b related data.

[0082] For example, by using causal conjunctions and syntactic analysis, the causal relationship between b second pieces of data can be determined, and the causal events can be obtained as b related data. For example, related data may include: interrelated technical means, technical problems, and application scenarios.

[0083] Step S503: Perform data splicing on the b associated data to generate c third data.

[0084] Figure 6 The flowchart illustrating the data generation method according to an embodiment of the present disclosure is shown.

[0085] like Figure 6 As shown, the method 600 includes steps S601 to S604.

[0086] Step S604: Merge the associated data with the same technical means in the b associated data to generate f first sub-associated data with different technical means. Each first sub-associated data includes the same technical means and all the corresponding technical problems and application scenarios. Here, f is greater than or equal to 0, f is less than or equal to b and f is an integer.

[0087] Step S602: Merge the related data of the same technical problem in the b related data to generate g second sub-related data of different technical problems. Each first sub-related data includes the same technical problem and all the corresponding technical means and application scenarios. g is greater than or equal to 0, g is less than or equal to b and g is an integer.

[0088] Step S603: Merge the associated data with the same application scenario in the b associated data to generate h third sub-associated data with different application scenarios. Each third sub-associated data includes the same application scenario and all the corresponding technical means and technical problems. Here, h is greater than or equal to 0, h is less than or equal to b and h is an integer.

[0089] Step S604: Merge and deduplicate the first sub-association data of f different technical means, the second sub-association data of g different technical problems, and the third sub-association data of h different application scenarios to generate c third data, wherein the sum of f, g, and h is greater than or equal to c.

[0090] By constructing data across three interconnected dimensions—technical methods, technical problems, and application scenarios—we generate technical problems and application scenarios under merged technical methods, technical methods and application scenarios under merged technical problems, and technical problems and technical methods under merged application scenarios. Merging and deduplicating these three elements improves the comprehensiveness of the coverage of c third-party data sets and enhances their reliability. Furthermore, by using sentence segmentation to generate causal conjunctions and grammatical logic, we not only improve the accuracy of the third-party data sets but also reduce manual intervention, saving time and manpower costs and increasing efficiency.

[0091] Return to reference Figure 2 In step S205, historical business data is obtained, and the historical business data is matched with c third data to generate a matching score for the c third data.

[0092] Figure 7 A flowchart illustrating the generation of a matching score for third data in a data generation method according to an embodiment of the present disclosure is shown.

[0093] like Figure 7 As shown, the method 700 includes steps S701 to S703.

[0094] Step S701: Obtain historical business data, extract semantic feature data from the historical business data, and generate a first semantic feature.

[0095] Step S702: Extract semantic features from c third data points to generate second semantic features for c third data points.

[0096] Step S703: Perform similarity matching scoring on the first semantic feature and the second semantic features of c third data respectively to generate matching scores for c third data.

[0097] For example, the cosine similarity algorithm can be used to calculate the similarity matching score between the first semantic feature and the second semantic features of c third data.

[0098] Return to reference Figure 2 In step S206, the matching score and dimension score of each of the c third data are weighted and calculated to generate a feasibility score for the c third data.

[0099] By generating matching scores for c third-party data based on historical business data, the usability and reliability of these c third-party data relative to the business can be improved, thereby increasing the efficiency and accuracy of subsequent decision-making.

[0100] Step S207: Obtain d third data whose feasibility scores are higher than the second preset threshold from c third data, and generate d target data, where d is less than or equal to c, d is greater than 0 and d is an integer.

[0101] After generating d target data points, a feasibility report can be generated and sent based on the d target data points.

[0102] The generated feasibility report can be integrated with existing enterprise R&D management, strategic planning, and other systems via standard API interfaces. The system provides both web and mobile application interfaces, supporting access and use from multiple terminals.

[0103] By generating and sending feasibility reports, data visualization and timeliness can be achieved, which facilitates further improvement in the efficiency and accuracy of subsequent decision-making.

[0104] Figure 8 A schematic block diagram of a data generation apparatus according to an embodiment of the present disclosure is shown.

[0105] like Figure 8 As shown, the device 800 includes: a first acquisition module 801, a first generation module 802, a second generation module 803, a third generation module 804, a fourth generation module 805, a fifth generation module 806, and a sixth acquisition module 807.

[0106] The first acquisition module 801 is used to acquire a first data points containing preset index terms from m data sources, where a is greater than 1 and is an integer, and m is greater than 1 and is an integer. In one embodiment, the first acquisition module 801 can be used to execute step S201 described above.

[0107] The first acquisition module 801 includes: a second acquisition module, a seventh generation module, a third acquisition module, and an eighth generation module.

[0108] The second acquisition module is used to acquire m retrieval rules from m data sources. In one embodiment, the second acquisition module can be used to execute step S301 described above, which will not be repeated here.

[0109] The seventh generation module is used to perform rule-based processing on the preset index terms based on the m retrieval rules, generating n retrieval expressions corresponding to the m data sources, where n is greater than or equal to m and n is an integer. In one embodiment, the seventh generation module can be used to execute step S302 described above, which will not be repeated here.

[0110] The third acquisition module is used to acquire e first data items containing preset index terms from m data sources based on the n search terms. In one embodiment, the third acquisition module can be used to execute step S303 described above, which will not be repeated here.

[0111] The eighth generation module is used to preprocess the e first data to generate a first data, where e is greater than or equal to a, and e is an integer. In one embodiment, the eighth generation module can be used to execute step S304 described above, which will not be repeated here.

[0112] The first generation module 802 is used to perform multi-dimensional scoring on the a first data points and generate a first data dimensional scores. In one embodiment, the first generation module 802 can be used to execute step S202 described above.

[0113] The first generation module 802 includes: a ninth generation module, a first calculation module, a tenth generation module, and an eleventh generation module.

[0114] The ninth generation module is used to extract text features from a first set of data and generate feature keywords for the a first set of data. In one embodiment, the ninth generation module can be used to execute step S401 described above, which will not be repeated here.

[0115] The first calculation module is used to calculate the inverse document frequency (IVF) of the a first data text based on the feature keywords of the a first data. In one embodiment, the first calculation module can be used to perform step S402 described above, which will not be repeated here.

[0116] The tenth generation module is used to perform trend prediction on the a first data points based on a time recurrent neural network model, and generate trend prediction values ​​for the a first data points. In one embodiment, the tenth generation module can be used to execute step S403 described above, which will not be repeated here.

[0117] The eleventh generation module is used to perform a weighted calculation of the word frequency inverse document frequency and the corresponding trend prediction value for each of the a first data points to generate a first data dimension scores. In one embodiment, the eleventh generation module can be used to execute step S404 described above, which will not be repeated here.

[0118] The second generation module 803 is used to acquire first data with scores of a first data dimension that are higher than a first preset threshold, and generate b second data, wherein b is less than or equal to a, b is greater than 0 and b is an integer. In one embodiment, the second generation module 803 can be used to execute step S203 described above, which will not be repeated here.

[0119] The third generation module 804 is used to perform causal analysis on b second data points to generate c third data points, where c is greater than 0 and c is an integer. In one embodiment, the third generation module 804 can be used to execute step S204 described above.

[0120] The third generation module 804 includes: the twelfth generation module, the thirteenth generation module, and the fourteenth generation module.

[0121] The twelfth generation module is used to perform sentence segmentation processing on each of the b second data, generating causal conjunctions and grammatical logic for the b second data. In one embodiment, the twelfth generation module can be used to execute step S501 described above, which will not be repeated here.

[0122] The thirteenth generation module is used to generate b related data based on the causal conjunctions and syntactic logic of the b second data. In one embodiment, the thirteenth generation module can be used to execute step S502 described above, which will not be repeated here.

[0123] The fourteenth generation module is used to concatenate the b associated data to generate c third data. In one embodiment, the fourteenth generation module can be used to execute step S503 described above.

[0124] The fourteenth generation module includes: the fifteenth generation module, the sixteenth generation module, the seventeenth generation module, and the eighteenth generation module.

[0125] The fifteenth generation module is used to merge the associated data with the same technical means among the b associated data to generate f first sub-associated data with different technical means. Each first sub-associated data includes the same technical means and all corresponding technical problems and application scenarios, where f is greater than or equal to 0, f is less than or equal to b, and f is an integer. In one embodiment, the fifteenth generation module can be used to execute step S601 described above, which will not be repeated here.

[0126] The sixteenth generation module is used to merge the related data of the same technical problem in the b related data to generate g second sub-related data of different technical problems. Each first sub-related data includes the same technical problem and all corresponding technical means and application scenarios, where g is greater than or equal to 0, g is less than or equal to b, and g is an integer. In one embodiment, the sixteenth generation module can be used to execute step S602 described above, which will not be repeated here.

[0127] The seventeenth generation module is used to merge the associated data with the same application scenario among the b associated data to generate h third sub-associated data with different application scenarios. Each third sub-associated data includes the same application scenario and all corresponding technical means and technical problems. Here, h is greater than or equal to 0, h is less than or equal to b, and h is an integer. In one embodiment, the seventeenth generation module can be used to execute step S603 described above, which will not be repeated here.

[0128] The eighteenth generation module is used to merge and deduplicate the first sub-association data of f different technical means, the second sub-association data of g different technical problems, and the third sub-association data of h different application scenarios to generate c third data, wherein the sum of f, g, and h is greater than or equal to c. In one embodiment, the eighteenth generation module can be used to execute step S604 described above, which will not be repeated here.

[0129] The fourth generation module 805 is used to acquire historical business data, match the historical business data with c third data points, and generate a matching score for the c third data points. In one embodiment, the fourth generation module 805 can be used to execute step S205 described above.

[0130] The fourth generation module 805 includes: the nineteenth generation module, the twentieth generation module, and the twenty-first generation module.

[0131] The nineteenth generation module is used to acquire historical business data, extract semantic feature data from the historical business data, and generate a first semantic feature. In one embodiment, the nineteenth generation module can be used to execute step S701 described above, which will not be repeated here.

[0132] The twentieth generation module is used to extract semantic features from c third data points and generate second semantic features for the c third data points. In one embodiment, the twentieth generation module can be used to execute step S702 described above, which will not be repeated here.

[0133] The twenty-first generation module is used to perform similarity matching and scoring on the first semantic feature and the second semantic features of c third data respectively, and generate matching scores for c third data. In one embodiment, the twenty-first generation module can be used to execute step S703 described above, which will not be repeated here.

[0134] The fifth generation module 806 is used to perform a weighted calculation of the matching degree score and dimension score of each of the c third data points to generate a feasibility score for the c third data points. In one embodiment, the fifth generation module 806 can be used to execute step S206 described above, which will not be repeated here.

[0135] The sixth acquisition module 807 is used to acquire d third data points whose feasibility scores are higher than the second preset threshold from c third data points, and generate d target data points, where d is less than or equal to c, d is greater than 0, and d is an integer. In one embodiment, the sixth acquisition module 807 can be used to execute step S207 described above, which will not be repeated here.

[0136] According to embodiments of this disclosure, any plurality of modules among the first acquisition module 801, the first generation module 802, the second generation module 803, the third generation module 804, the fourth generation module 805, the fifth generation module 806, and the sixth acquisition module 807 can be combined into one module, or any one of these modules can be split into multiple modules. Alternatively, at least a portion of the functionality of one or more of these modules can be combined with at least a portion of the functionality of other modules and implemented in one module. According to embodiments of this disclosure, at least one of the first acquisition module 801, the first generation module 802, the second generation module 803, the third generation module 804, the fourth generation module 805, the fifth generation module 806, and the sixth acquisition module 807 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging the circuitry, or implemented in any one of software, hardware, and firmware methods, or in a suitable combination of any of these methods. Alternatively, at least one of the first acquisition module 801, the first generation module 802, the second generation module 803, the third generation module 804, the fourth generation module 805, the fifth generation module 806, and the sixth acquisition module 807 can be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.

[0137] Figure 9 A block diagram schematically illustrates an electronic device suitable for implementing a data generation method according to an embodiment of the present disclosure.

[0138] like Figure 9As shown, an electronic device 900 according to an embodiment of the present disclosure includes a processor 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage portion 908 into a random access memory (RAM) 903. The processor 901 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 901 may also include onboard memory for caching purposes. The processor 901 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.

[0139] RAM 903 stores various programs and data required for the operation of electronic device 900. Processor 901, ROM 902, and RAM 903 are interconnected via bus 904. Processor 901 performs various operations of the method flow according to embodiments of the present disclosure by executing programs in ROM 902 and / or RAM 903. It should be noted that the programs may also be stored in one or more memories other than ROM 902 and RAM 903. Processor 901 may also perform various operations of the method flow according to embodiments of the present disclosure by executing programs stored in said one or more memories.

[0140] According to embodiments of this disclosure, the electronic device 900 may further include an input / output (I / O) interface 905, which is also connected to a bus 904. The electronic device 900 may also include one or more of the following components connected to the I / O interface 905: an input section 906 including a keyboard, mouse, etc.; an output section 907 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a LAN card, modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the I / O interface 905 as needed. A removable medium 911, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 910 as needed so that computer programs read from it can be installed into the storage section 908 as needed.

[0141] This disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs that, when executed, implement the method according to the embodiments of this disclosure.

[0142] According to embodiments of this disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, such as including, but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this disclosure, the computer-readable storage medium may include ROM 902 and / or RAM 903 and / or one or more memories other than ROM 902 and RAM 903 described above.

[0143] Embodiments of this disclosure also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code is used to cause the computer system to implement the data generation method provided in the embodiments of this disclosure.

[0144] When the computer program is executed by the processor 901, it performs the functions defined in the system / apparatus of this disclosure embodiments. According to embodiments of this disclosure, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0145] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and downloaded and installed via the communication section 909, and / or installed from a removable medium 911. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.

[0146] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 909, and / or installed from the removable medium 911. When the computer program is executed by the processor 901, it performs the functions defined in the system of this disclosure embodiment. According to embodiments of this disclosure, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0147] According to embodiments of this disclosure, program code for executing the computer programs provided in embodiments of this disclosure can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C", or similar programming languages. The program code can execute entirely on a user's computing device, partially on a user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0148] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0149] Those skilled in the art will understand that the features described in the various embodiments and / or claims of this disclosure can be combined or combined in various ways, even if such combinations or combinations are not explicitly described in this disclosure. In particular, the features described in the various embodiments and / or claims of this disclosure can be combined or combined in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or combinations fall within the scope of this disclosure.

[0150] The embodiments of this disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of this disclosure. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. The scope of this disclosure is defined by the appended claims and their equivalents. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of this disclosure, and all such substitutions and modifications should fall within the scope of this disclosure.

Claims

1. A data generation method, characterized in that, The method includes: Retrieve a first data points containing preset index terms from m data sources, where a is greater than 1 and a is an integer, and m is greater than 1 and m is an integer; Perform multi-dimensional scoring on the a first data points to generate a first data point dimension scores; Obtain first data with scores of a first data dimension that are higher than the first preset threshold, and generate b second data, where b is less than or equal to a, b is greater than 0 and b is an integer; Perform causal analysis on b second data points to generate c third data points, where c is greater than 0 and c is an integer; Obtain historical business data, match the historical business data with c third data, and generate a matching score for the c third data; A weighted average of the matching score and dimension score for each of the c third-data points is calculated to generate a feasibility score for the c third-data points; and Obtain c third data points whose feasibility scores are higher than the second preset threshold, and generate d target data points, where d is less than or equal to c, d is greater than 0 and d is an integer.

2. The method according to claim 1, characterized in that, Retrieve a first data points containing preset index terms from m data sources, including: Obtain m search rules from m data sources; Based on the m retrieval rules, the preset index terms are processed to generate n retrieval formulas corresponding to the m data sources, where n is greater than or equal to m and n is an integer; Based on the n search terms, obtain e first data points containing preset index terms from m data sources; and The e first data are preprocessed to generate a first data, where e is greater than or equal to a and e is an integer.

3. The method according to claim 1, characterized in that, The a first data points are scored in multiple dimensions to generate a first data point dimension scores, including: Extract text features from a set of a first data points to generate feature keywords for the a first data points; Based on the feature keywords of the a first data, calculate the term frequency inverse document frequency of the a first data text; Based on a time recurrent neural network model, trend prediction is performed on the *a* first data points to generate trend prediction values ​​for the *a* first data points; and The word frequency inverse document frequency and the corresponding trend prediction value of each of the a first data points are weighted and calculated to generate a first data dimension scores.

4. The method according to claim 1, characterized in that, Perform causal analysis on b second data points to generate c third data points, including: Sentence segmentation is performed on each of the b second data to generate causal conjunctions and grammatical logic for the b second data; Based on the causal conjunctions and grammatical logic of the b second data, generate b related data; and The b related data are concatenated to generate c third data.

5. The method according to claim 4, characterized in that, The associated data includes: interrelated technical means, technical problems, and application scenarios. The b pieces of associated data are concatenated to generate c pieces of third-party data, including: The data with the same technical means in the b associated data are merged to generate f first sub-associated data with different technical means. Each first sub-associated data includes the same technical means and all the corresponding technical problems and application scenarios. Here, f is greater than or equal to 0, f is less than or equal to b and f is an integer. The data related to the same technical problem in the b related data are merged to generate g second sub-related data with different technical problems. Each first sub-related data includes the same technical problem and all corresponding technical means and application scenarios. g is greater than or equal to 0, g is less than or equal to b and g is an integer. The b associated data sets are merged to generate h third-level sub-associated data sets with different application scenarios. Each third-level sub-associated data set includes the same application scenario and all corresponding technical means and technical problems. Here, h is greater than or equal to 0, h is less than or equal to b, and h is an integer. The first sub-association data of f different technical means, the second sub-association data of g different technical problems, and the third sub-association data of h different application scenarios are merged and deduplicated to generate c third data, wherein the sum of f, g, and h is greater than or equal to c.

6. The method according to any one of claims 1 to 5, characterized in that, Obtain historical business data, match the historical business data with c third-party data, and generate a matching score for the c third-party data, including: Acquire historical business data, extract semantic feature data from the historical business data, and generate a first semantic feature; Semantic features are extracted from c third data points to generate second semantic features for c third data points; The first semantic feature is matched and scored with the second semantic features of c third data respectively to generate matching scores for c third data.

7. The method according to any one of claims 1 to 5, characterized in that, The method also includes: Based on the d target data, a feasibility report is generated and sent.

8. A data generation apparatus, characterized in that, The device includes: The first acquisition module is used to acquire a first data points containing preset index terms from m data sources, where a is greater than 1 and a is an integer, and m is greater than 1 and m is an integer; The first generation module is used to perform multi-dimensional scoring on the a first data points and generate a first data dimensional scores. The second generation module is used to obtain first data with scores of a first data dimension that are higher than the first preset threshold, and generate b second data, wherein b is less than or equal to a, b is greater than 0 and b is an integer. The third generation module is used to perform causal analysis on b second data points and generate c third data points, where c is greater than 0 and c is an integer. The fourth generation module is used to acquire historical business data, match the historical business data with c third data, and generate a matching score for the c third data. The fifth generation module is used to perform a weighted calculation of the matching degree score and dimension score of each of the c third data points to generate a feasibility score for the c third data points; and The sixth generation module is used to obtain d third data whose feasibility scores of c third data are higher than the second preset threshold, and generate d target data, where d is less than or equal to c, d is greater than 0 and d is an integer.

9. An electronic device, comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 7.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.