A blockchain data collection method and device based on a multi-architecture NLP pre-training model
By combining multi-architecture NLP pre-trained models with blockchain technology, this data collection method solves the problem of manual reliance in enterprise surveys and evaluations, improves data processing efficiency and accuracy, and achieves intelligent support and credibility of data.
Patent Information
- Application Number
- CN202310612594.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-29
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2043-05-29
AI Technical Summary
In business environment and industrial park services consulting and other research and evaluation services, enterprises rely on personnel experience and labor-intensive methods for research design and data result evaluation. Traditional methods cannot replace manual processing on a large scale, resulting in low processing efficiency and inaccurate data.
A blockchain data acquisition method based on multi-architecture NLP pre-trained models is adopted. Two pre-trained models are used to process preset demand information and market research data respectively to generate design schemes and data results. The consensus mechanism of blockchain is used to ensure the credibility and immutability of the data.
It improved data processing efficiency and accuracy, reduced manual input, ensured data credibility and integrity, and achieved intelligent data support.
Smart Images

Figure CN116701930B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the technical field of artificial intelligence, in particular to a multi-architecture NLP pre-training model and a data collection and evaluation method and device based on a blockchain. BACKGROUND
[0002] With the continuous development of big data business, enterprises accumulate massive multi-dimensional big data in the long-term business expansion process. However, the application of these data is mainly limited to data archiving and data querying, and has not yet fully tapped its data value through knowledge distillation, knowledge graph, knowledge data, and data intelligence. These "sleeping" big data cannot provide intelligent support for data processing in the research and analysis stage, significantly reducing manual input and effectively solving data collection problems.
[0003] In the research and evaluation business of business environment and park services for enterprise consultation, research scheme design and data result evaluation are more dependent on personnel experience and labor-intensive processing links. Traditional research scheme design mainly relies on manual processing, and the intelligent level of information-based auxiliary data processing is relatively limited. Currently, it mainly focuses on voice text translation, data query and retrieval, and other primary levels, and cannot replace manual work to a large extent. In the design of scheme data and data conclusion data for collecting information, manual processing of data often leads to low processing efficiency and inaccurate data. SUMMARY
[0004] The purpose of the present application is to provide a blockchain data collection method based on a multi-architecture NLP pre-training model, which can improve data processing efficiency and data accuracy.
[0005] In a first aspect, the application provides a blockchain data collection method based on a multi-architecture NLP pre-training model, which adopts the following technical solution:
[0006] A blockchain data collection method based on a multi-architecture NLP pre-training model, comprising:
[0007] In response to a request to obtain an initial training data set, and training different architecture models according to the initial training data set to obtain a first pre-training model and a second pre-training model;
[0008] Training the first pre-training model and the second pre-training model according to the preset demand information to obtain a first design scheme and a second design scheme;
[0009] According to the first preset rule, the first design scheme and the second design scheme are analyzed to obtain a first analysis result, and a first positive sample data set of the first design scheme and the second design scheme is generated according to the first analysis result, and a final design scheme is obtained according to the first positive sample data set, and market research data is collected according to the final design scheme;
[0010] According to the first pre-training model and the second pre-training model, the market research data is analyzed to obtain a first preliminary data result and a second preliminary data result;
[0011] According to the second preset rule, the first preliminary data result and the second preliminary data result are analyzed to obtain a second analysis result, and a second positive sample data set of the first preliminary data result and the second preliminary data result is generated according to the second analysis result, and a final data result is obtained according to the second positive sample data set;
[0012] According to the first positive sample data set and the second positive sample data set, new data content in the blockchain is determined, and a new block is generated according to the new data content, and the new block is responded and authenticated.
[0013] By adopting the above technical scheme, the preset demand information is processed by two pre-training models respectively, two design schemes can be obtained, and then the first positive sample data set in the two design schemes is screened out according to the first preset rule, and the final design scheme is generated according to the first positive sample data set, so that the data processing efficiency and data accuracy in the scheme design stage can be improved. By analyzing and processing the market research data by two pre-training models, two data results can be obtained, and then the second positive sample data set in the two data results is screened out according to the second preset rule, and the final data result is generated according to the second positive sample data set, so that the data processing efficiency and data accuracy in the data conclusion stage can be improved.
[0014] Optionally, the initial training data set is obtained in response to the request, and the model of different architectures is trained according to the initial training data set to obtain the first pre-training model and the second pre-training model, specifically including:
[0015] An open data set is obtained, and the market research special sample data set is generated according to the market research related data, and the open data set and the market research special sample data set are fused into the initial training data set;
[0016] training the GPT model of the Decoder structure of the Transformer by calling the initial training data set to obtain the first pre-training model, and testing the first pre-training model;
[0017] training the BERT model of the Encoder structure of the Transformer by calling the initial training data set to obtain the second pre-training model, and testing the second pre-training model.
[0018] By using the above technical solutions, the pre-training models of two structures can be generated by training the models of two structures respectively by using the generated initial training data set, so that the pre-training models of two different structures are conveniently called for subsequent processing of data information.
[0019] Optionally, the first design scheme and the second design scheme are obtained by training the preset requirement information according to the first pre-training model and the second pre-training model, specifically including:
[0020] processing the preset requirement information by calling the first pre-training model to obtain a first whole-process workflow, and defining the input, process and output of each link in the first whole-process workflow according to a preset first input module library, a first process module library and a first output module library;
[0021] processing and analyzing the input, process and output of each link in the first whole-process workflow by calling the first pre-training model to obtain first input data, first process data and first output data of each link, and fusing the first input data, the first process data and the first output data to obtain the first design scheme;
[0022] processing and analyzing the input, process and output of each link in the first whole-process workflow by calling the second pre-training model to obtain second input data, second process data and second output data of each link, and fusing the second input data, the second process data and the second output data to obtain the second design scheme.
[0023] By using the above technical solutions, the first whole-process workflow can be obtained by processing the preset requirement information by calling the first pre-training model, and then the input, process and data of each link in the first whole-process workflow are defined according to the preset module library, and then the defined input, process and output are analyzed and processed by calling the first pre-training model and the second pre-training model respectively to obtain the data of each link, and the obtained data of each link are fused to generate the first design scheme and the second design scheme respectively, so as to improve the accuracy of data processing and thus improve the accuracy of the generated design scheme.
[0024] Optionally, the first design scheme and the second design scheme are analyzed according to a first preset rule to obtain a first analysis result, the first design scheme and the second design scheme are generated into a first positive sample data set according to the first analysis result, a final design scheme is obtained according to the first positive sample data set, and market research data is collected according to the final design scheme, specifically including:
[0025] The second pre-trained model is called to evaluate the consistency of the first design scheme and the second design scheme, and a first consistency evaluation result is obtained.
[0026] The first consistency evaluation result is quantitatively analyzed to obtain a first overall quantitative evaluation index of the first design scheme and the second design scheme, and if the first overall quantitative evaluation index exceeds a first preset threshold, a credibility evaluation of the first design scheme and the second design scheme is performed to obtain a first credibility evaluation result.
[0027] The first design scheme and the second design scheme are processed according to the first credibility evaluation result to obtain a plurality of first positive sample data that can be adopted.
[0028] The first positive sample data is fused to obtain the first positive sample data set, the first positive sample data set is processed to obtain the final design scheme, market research data is collected according to the final design scheme, and the market research data is associated with a newly created market research task.
[0029] By using the above technical solution, the pre-trained models of two structures are called to analyze and process the first design scheme and the second design scheme, and the second pre-trained model is called to check the first design scheme and the second design scheme, so that the final design scheme can be generated, avoiding the limitation of data processing by a single structure model, thereby improving the data accuracy in the scheme design stage.
[0030] Optionally, the market research data is analyzed according to the first pre-trained model and the second pre-trained model to obtain first preliminary data results and second preliminary data results, specifically including:
[0031] The first pre-trained model is called to process the market research data to obtain a second whole-process workflow, and a second input module library, a second process module library and a second output module library are defined according to a preset to define the input, process and output of each link in the second whole-process workflow.
[0032] The first pre-trained model is invoked to process and analyze the inputs, processes, and outputs of each stage in the second full-process workflow, to obtain the third input data, third process data, and third output data of each stage, and the third input data, the third process data, and the third output data are fused to obtain the first preliminary data result;
[0033] The second pre-trained model is invoked to process and analyze the inputs, processes, and outputs of each stage in the second full-process workflow, to obtain the fourth input data, fourth process data, and fourth output data of each stage. The fourth input data, the fourth process data, and the fourth output data are then fused to obtain the second preliminary data results.
[0034] By adopting the above technical solution, the first pre-trained model is called to process the market research data to obtain the second full-process workflow. Then, according to the preset module library, the inputs, processes and data of each link in the second full-process workflow are defined. The first and second pre-trained models are then called to analyze and process the defined inputs, processes and outputs to obtain the data of each link. The data of each link are then fused to generate the first preliminary data results and the second preliminary data results, thereby improving the accuracy of data processing and thus improving the accuracy of the generated data conclusions.
[0035] Optionally, the step of analyzing the first preliminary data result and the second preliminary data result according to a second preset rule to obtain a second analysis result, generating a second positive sample dataset of the first preliminary data result and the second preliminary data result based on the second analysis result, and obtaining the final data result based on the second positive sample dataset specifically includes:
[0036] The second pre-trained model is invoked to evaluate the consistency between the first preliminary data result and the second preliminary data result, and a second consistency evaluation result is obtained.
[0037] The second consistency evaluation result is quantitatively analyzed to obtain a second overall quantitative evaluation index of the first preliminary data result and the second preliminary data result. If the second overall quantitative evaluation index exceeds a second preset threshold, the credibility of the first preliminary data result and the second preliminary data result is evaluated to obtain a second credibility evaluation result.
[0038] Based on the second credibility evaluation result, the first preliminary data result and the second preliminary data result are processed to obtain multiple credible second positive sample data;
[0039] Fuse a plurality of second positive sample data that can be accepted, obtain the second positive sample data set, and process the second positive sample data set to obtain the final data result.
[0040] By adopting the technical scheme, the two kinds of pre-training models are called to analyze and process the first preliminary data result and the second preliminary data result, and the second pre-training model is called to check the first preliminary data result and the second preliminary data result, so that the final data result can be generated, the limitation of a single structure model on data processing is avoided, and the data accuracy in the data result stage can be improved.
[0041] Optionally, the new data content in the blockchain is determined according to the first positive sample data set and the second positive sample data set, and a new block is generated according to the new data content, specifically including:
[0042] The distributed node obtains authorization permission from the CA root node server;
[0043] If the distributed node obtains the first positive sample data set and the second positive sample data set with credibility, a new block is automatically generated by the distributed node, the new block is encrypted by using the key stored by the distributed node, and the new block is broadcasted to other distributed nodes.
[0044] By adopting the technical scheme, when the distributed node obtains the first positive sample data set and the second positive sample data set with credibility, a new block can be automatically generated by the distributed node, and then the new block is broadcasted to other distributed nodes, so that the new block can be conveniently generated according to the first positive sample data set and the second positive sample data set.
[0045] Optionally, the new block is responded and authenticated, specifically including:
[0046] The other distributed nodes of the private chain of the blockchain authenticate the new block according to the consensus mechanism of the private chain, and after the authentication is passed, the new block takes effect, and the first positive sample data set and the second positive sample data set are automatically linked to the tail of the private chain;
[0047] When the new block takes effect, the CA root node server re-packages the data content of the new block, encrypts the new block by using the password stored by the CA root node server, and broadcasts the new block to other distributed nodes of the public chain of the blockchain;
[0048] Other distributed nodes of the public chain authenticate the new block according to the consensus mechanism of the public chain, and after the authentication is passed, the new block takes effect, and the first positive sample data set and the second positive sample data set are automatically linked to the tail of the public chain.
[0049] By adopting the technical solution, the distributed ledger technology of the block chain is adopted, the new block is generated according to the positive sample data set after consistency analysis and credibility analysis, the new block is authenticated based on the consensus mechanism of the block chain, the new block takes effect after the authentication is passed, and the positive sample data set is automatically linked to the block chain, so that the credibility and non-tamperability of the design scheme data and the data result data are ensured.
[0050] Optionally, a first negative sample data set of the first design scheme and the second design scheme is generated according to the first analysis result, a second negative sample data set of the first preliminary data result and the second preliminary data result is generated according to the second analysis result, and the data content of the new block is fused into the first positive sample data set and the second positive sample data set respectively, and the first positive sample data set, the second positive sample data set, the first negative sample data set and the second negative sample data set are fused into the market research special sample data set.
[0051] By adopting the technical solution, the first positive sample data set, the second positive sample data set, the first negative sample data set and the second negative sample data set are fused into the market research special sample data set, and the market research special sample data set and the open data set are fused again to obtain a new initial training data set, and the new initial training data set is used to train the models of the two architectures respectively, so that new first pre-training models and second pre-training models can be obtained to improve the model training performance of the first pre-training models and the second pre-training models.
[0052] In a second aspect, the application provides a blockchain data acquisition method based on a multi-architecture NLP pre-training model, which adopts the following technical solution:
[0053] A blockchain data acquisition method based on a multi-architecture NLP pre-training model, comprising:
[0054] A model construction module is configured to obtain an initial training data set in response to a request, and train models of different architectures according to the initial training data set to obtain first pre-training models and second pre-training models.
[0055] A first generation module is configured to train preset demand information according to the first pre-training models and the second pre-training models to obtain first design schemes and second design schemes.
[0056] The first analysis module is configured to analyze the first design scheme and the second design scheme according to a first preset rule, obtain a first analysis result, generate a first positive sample data set of the first design scheme and the second design scheme according to the first analysis result, obtain a final design scheme according to the first positive sample data set, and collect market research data according to the final design scheme.
[0057] The second generation module is configured to analyze the market research data according to the first pre-training model and the second pre-training model respectively, and obtain a first preliminary data result and a second preliminary data result.
[0058] The second analysis module is configured to analyze the first preliminary data result and the second preliminary data result according to a second preset rule, obtain a second analysis result, generate a second positive sample data set of the first preliminary data result and the second preliminary data result according to the second analysis result, and obtain a final data result according to the second positive sample data set.
[0059] The third generation module is configured to determine new data content in the blockchain according to the first positive sample data set and the second positive sample data set, generate a new block according to the new data content, and respond to and authenticate the new block.
[0060] By using the above technical solution, the two pre-training models are used to process the preset demand information respectively, two design schemes are obtained, and then the first preset rule is used to check and analyze the two design schemes to screen out the first positive sample data set in the two design schemes, and the final design scheme is generated according to the first positive sample data set, so as to improve the data processing efficiency and data accuracy in the scheme design stage. The two pre-training models are used to analyze and process the market research data respectively, two data results are obtained, and then the second preset rule is used to check and analyze the two data results to screen out the second positive sample data set in the two data results, and the final data result is generated according to the second positive sample data set, so as to improve the data processing efficiency and data accuracy in the data conclusion stage. BRIEF DESCRIPTION OF DRAWINGS
[0061] Figure 1 is a method flowchart of steps S100-S600 in the embodiment of the present application;
[0062] Figure 2 is a method flowchart of steps S110-S130 in the embodiment of the present application;
[0063] Figure 3 is a method flowchart of steps S210-S230 in the embodiment of the present application;
[0064] Figure 4 is a method flowchart of steps S310-S340 in the embodiment of the present application;
[0065] Figure 5 is a method flowchart of steps S410-S430 in the embodiment of the present application;
[0066] Figure 6 is a method flowchart of steps S510-S540 in the embodiment of the present application;
[0067] Figure 7 is a method flowchart of steps S610-S620 in the embodiment of the present application;
[0068] Figure 8 is a method flowchart of steps S630-S650 in the embodiment of the present application;
[0069] Figure 9 is a module framework diagram of the structure and on-chain process of the public chain and the private chain in the embodiment of the present application. DETAILED DESCRIPTION
[0070] The following will be described in detail in combination with the accompanying drawings. Figure 1 - the accompanying drawings Figure 9 The present application will be further described in detail.
[0071] In order to fully tap the value of "sleeping" big data and make the expert experience explicit, it is necessary to apply the cutting-edge NLP model technology results to convert the traditional heavily dependent manual processing link into a "computable and quantifiable" problem. It is necessary to use the method of combining multi-architecture NLP pre-training model and blockchain technology to establish a new data collection and evaluation system of data intelligent processing and adoption, artificial assistance and key decision-making, and effectively solve the problem of data adoption relying on manual and data tampering in the market research process.
[0072] NLP pre-training model is a natural language processing tool driven by artificial intelligence technology. It is based on NLP algorithm model to jointly train massive public parameters and enterprise big data platform module for research analysis and data collection. It can understand and learn human language, especially in data collection, research analysis, and provide professional semantic analysis and output reasoning and judgment results. According to the user's agreed format, output text, language and charts.
[0073] NLP models with different technical structures refer to NLP models with different technical routes or pre-training methods. NLP pre-training models for preprocessing focus more on semantic understanding, information synthesis, reasoning ability and model generalization ability. NLP pre-training models for secondary review highlight data classification, entity extraction, sentiment judgment and accurate prediction ability.
[0074] WBS (Work Breakdown Structure) and factor decomposition is a principle that divides a project according to certain principles, divides the project into tasks, and further divides the tasks into a series of work, and then allocates a series of work to everyone's daily activities until it cannot be further divided, that is, project-task-work-daily activity.
[0075] A blockchain data collection method based on a multi-architecture NLP pre-training model, referring to Figure 1 , specifically comprising the following steps:
[0076] S100: In response to a request, an initial training data set is obtained, and different architecture models are trained according to the initial training data set to obtain a first pre-training model and a second pre-training model.
[0077] In an embodiment of the present application, referring to Figure 2 , step S100 specifically comprises the following steps:
[0078] S110: Obtain an open data set, generate a market research special sample data set according to market research related data, and fuse the open data set and the market research special sample data set into an initial training data set.
[0079] In this embodiment, the open data set is an open source natural language processing model training data set, and a market research special sample data set is generated according to the market research big data accumulated by the enterprise big data center, and the open data set and the market research special sample data set are fused to obtain an initial training data set.
[0080] S120: Train the GPT model of the Decoder structure of the Transformer by calling the initial training data set, obtain the first pre-training model, and test the first pre-training model.
[0081] S130: Train the BERT model of the Encoder structure of the Transformer by calling the initial training data set, obtain the second pre-training model, and test the second pre-training model.
[0082] Wherein, the data set is used to train the GPT model of the Decoder structure of the Transformer and the BERT model of the Encoder structure of the Transformer to obtain the pre-training model, which is a conventional technique in the art; wherein, in an embodiment of the present application, the first pre-training model and the second pre-training model also need to be tested and deployed according to the timing or pre-set rules to improve the accuracy of the first pre-training model and the second pre-training model, facilitating subsequent use; wherein, in this embodiment, the training method of the two architecture models refers to the training method of large models such as ChatGPT.
[0083] S200: training the preset demand information according to the first pre-trained model and the second pre-trained model respectively to obtain the first design scheme and the second design scheme.
[0084] In an embodiment of the present application, with reference to Figure 3 , step S200 specifically includes the following steps:
[0085] S210: calling the first pre-trained model to process the preset demand information to obtain a first whole-process workflow, and defining the input, process and output of each link in the first whole-process workflow according to the preset first input module library, first process module library and first output module library.
[0086] In an embodiment of the present application, the preset demand information can be a market overall research and analysis demand, for example, the preset demand information is a new market research task, and the overall research and analysis demand of the specific market research task is established in a computer-aided manner to describe the overall requirement of the market research task. Specifically, the form of the overall research and analysis demand can include voice, text, video and picture, etc.
[0087] Specifically, in the present embodiment, the acquisition method of the overall research and analysis demand specifically includes the following steps:
[0088] Obtain new market research related files, and the new market research related files include market research project tasks, task overall demand and task description; classify the new market research task according to the new market research related files to obtain classification types of each dimension, determine attribute information and category information of the new market research task according to the classification types, and construct the overall research and analysis demand according to the attribute information and the category information.
[0089] In the present embodiment, the open source NLP large model dataset is mainly used to screen out the content related to market research, and in addition, the cases, reports and data accumulated by enterprises in historical business are reorganized according to the format of the large model dataset to supplement the market research special sample dataset.
[0090] Specifically, the cases, reports and data accumulated by the enterprise in historical business include market research process and result files and data such as research questionnaires and databases, interview results, reports, etc. In the embodiment, a data collection and evaluation information service system meeting the market research full life cycle management is established, and a new market research project task is created based on the system, the overall demand and task description of the new task are set, and research and analysis demand files such as voice, text, pictures and videos are imported. The data collection and evaluation information service system for market research full life cycle management in the embodiment is a commonly used system in the field, so it will not be described here. With the support of the big data center, users can carry out data management, research plan, data entry and analysis operations, and can realize the coverage of the whole market research process.
[0091] More specifically, in the embodiment, the classification types of each dimension of the new market research task are defined, the attributes and classification information of the new task are determined, and the research and analysis task creation based on the overall demand is finally completed.
[0092] For example, for a city-level business environment evaluation task, the geographical location (such as a province, a city, etc.), the service type (such as business environment, etc.), the service object (such as government agencies, etc.), the evaluation model (such as the World Bank business environment evaluation system, etc.), the scheme design stage (such as scheme design and user research, etc.), the evaluation stage (such as data evaluation and report preparation, etc.) of the city should be determined respectively.
[0093] Among them, the data collection and evaluation information service system is established to realize the management support for the whole process from research scheme design to data result evaluation of market research; and based on the system, the research and analysis requirements of the specific market research task are established to describe the overall requirements of the market research task.
[0094] Among them, the first pre-training model is called to analyze and process the market overall research and analysis requirements, and the first whole process workflow can be obtained. The first whole process workflow in the embodiment is the whole process workflow of the market research scheme design stage, and the input, process and output of each link in the first whole process workflow can be defined according to the first input module library, the first process module library and the first output module library in the above data collection and evaluation information service system, to complete the creation and process decomposition of the task based on the market overall research and analysis requirements. Specifically, the definition process is determined by the user through human-computer interaction, mainly based on high-level requirements and artificial experience.
[0095] The first whole-process workflow is a work package decomposed by the data collection and evaluation information service system according to the business process and recombined into a workflow according to the business flow mode; and each work package contains input, process and output, for example, the questionnaire design work package, the input is the above-mentioned determined attribute information, classification information and raw corpus data, the process can be AI intelligent design and expert experience, that is, the data collection and evaluation information service system calls the pre-trained model to automatically design the questionnaire, and at the same time, manual review based on expert experience is required, and the output can be the questionnaire design result, which is used as the input of the next link.
[0096] Among them, facing the business environment evaluation requirements, respectively for the task requirements of the four stages of scheme design, user research, data evaluation and report preparation, the minimum work package "input-process-output" model based on work breakdown structure (WBS) is established, and the first pre-trained model is used to process the minimum work package to obtain the input, process and output parameters or file template requirements of the minimum work package.
[0097] Among them, creating work breakdown structure (WBS) is the process of decomposing project deliverables and project work into smaller, more manageable components; the main role of this process is to provide a structured view of the content to be delivered.
[0098] For example, taking the business environment evaluation sub-dimension "access to electricity" as an example, based on the evaluation dimension setting of the World Bank, six secondary WBS work packages of "access to electricity link", "access to electricity time", "access to electricity cost", "power supply reliability and electricity fee transparency index", "electricity price" and "access to electricity convenience" are established for the first level WBS work package "access to electricity", and "input-process-output" model is established respectively.
[0099] S220: calling the first pre-trained model to process and analyze the input, process and output of each link in the first whole-process workflow, obtaining the first input data, first process data and first output data of each link, and fusing the first input data, first process data and first output data to obtain the first design scheme.
[0100] Among them, in this embodiment, the first input data is the first input design scheme, the first process data is the first process execution scheme, and the first output data is the first output result.
[0101] Among them, in this embodiment, the first pre-trained model is called to design the input, process and output of each link in the first whole-process workflow defined in step S210, and then generate the first input design scheme, the first process execution scheme and the first output result for each link.
[0102] In the embodiment, the first pre-trained model is called to form a scheme set including a research questionnaire, expert interview content and scheme, data list, process tool or function module authorization, and finally complete the data definition of the WBS work package "input-process-output" model.
[0103] The first input design scheme, the first process execution scheme and the first output result of each stage are fused to finally obtain the first design scheme.
[0104] The language understanding and information synthesis capability of the first pre-trained model is used to analyze and predict the overall research and analysis requirement, and then the first design scheme and the required data collection requirement can be generated, and the analysis and prediction results can be output in a preset file and format for next step processing.
[0105] S230: The second pre-trained model is called to process and analyze the input, process and output of each link in the first whole process workflow, to obtain the second input data, the second process data and the second output data of each link, and fuse the second input data, the second process data and the second output data to obtain the second design scheme.
[0106] In the embodiment, the second input data is the second input design scheme, the second process data is the second process execution scheme, and the second output data is the second output result.
[0107] In the embodiment, the method of calling the second pre-trained model to process and analyze the input, process and output of each link in the first whole process workflow is the same as the method of calling the first pre-trained model to process and analyze the input, process and output of each link in the workflow, which will not be described here.
[0108] The second input design scheme, the second process execution scheme and the second output result of each stage are fused to finally obtain the second design scheme.
[0109] In the embodiment, the inference capability of the second pre-trained model is used to analyze the overall research and analysis requirement, and then the second design scheme and the required data collection requirement can be generated, and the analysis results can be output in a preset file and format for next step processing.
[0110] S300: analyzing the first design scheme and the second design scheme according to a first preset rule to obtain a first analysis result, generating a first positive sample data set of the first design scheme and the second design scheme according to the first analysis result, obtaining a final design scheme according to the first positive sample data set, and collecting market research data according to the final design scheme.
[0111] In an embodiment of the present application, with reference to Figure 4 , step S300 specifically includes the following steps:
[0112] S310: calling the second pre-trained model to perform consistency evaluation on the first design scheme and the second design scheme to obtain a first consistency evaluation result.
[0113] In the embodiment, the second pre-trained model is called to perform consistency evaluation on the first design scheme obtained in step S220 and the second design scheme obtained in step S230 according to the input, process and output of each link in the first whole-process workflow, and a first consistency evaluation result is obtained, and then whether the first design scheme and the second design scheme meet the requirement of consistency is judged according to the first consistency evaluation result.
[0114] In the embodiment, the basis for consistency evaluation comes from the understanding and reasoning of the second pre-trained model on the overall research and analysis demand; and the result of consistency evaluation includes overall qualitative evaluation and preset multi-dimension consistency quantitative evaluation, and the second pre-trained model will give comparative processing results for inconsistent contents.
[0115] In the embodiment, according to the input, process and output of each link defined in the first whole-process workflow, the second pre-trained model is called to perform consistency evaluation on the first input design scheme and the second input design scheme, the first process execution scheme and the second process execution scheme, and the first output result and the second output result, respectively, to obtain a consistency evaluation result between the first input design scheme and the second input design scheme, a consistency evaluation result between the first process execution scheme and the second process execution scheme, and a consistency evaluation result between the first output result and the second output result.
[0116] In the embodiment, the second pre-trained model can not only analyze the preset demand information and output the second design scheme like the first pre-trained model, but also use the semantic understanding and logical reasoning function of the second pre-trained model itself to perform consistency evaluation on the first design scheme and the second design scheme.
[0117] S320: quantitatively analyze the first consistency evaluation result to obtain a first overall quantitative evaluation index of the first design scheme and the second design scheme, and if the first overall quantitative evaluation index exceeds a first preset threshold, perform a reliability evaluation on the first design scheme and the second design scheme to obtain a first reliability evaluation result.
[0118] In this embodiment, the first consistency evaluation result is quantitatively analyzed to obtain first quantitative evaluation indexes of the input, process and output of each link in the first design scheme and the second design scheme, and the first quantitative evaluation indexes are weighted and calculated according to a preset weight to obtain the first overall quantitative evaluation index of the first design scheme and the second design scheme.
[0119] In this embodiment, for the first quantitative evaluation index of each WBS work package in the business environment evaluation task, a hierarchical dimension weight matrix based on the WBS decomposition level is established using the AHP (AHP), and the first quantitative evaluation index of each sub-dimension is weighted and calculated to finally obtain the first overall quantitative evaluation index; the output result of each link is divided into a first dimension, a second dimension,..., and a bottom dimension, and each level of dimension includes a series of indexes, and a weight is set for each index.
[0120] If the first overall quantitative evaluation index exceeds the first preset threshold, the reliability of the first design scheme and the second design scheme is evaluated according to the input, process and output of each link in the first design scheme and the second design scheme, respectively, to obtain the first reliability evaluation result.
[0121] In this embodiment, the first preset threshold is set to 0.6 (after normalization, the full score is 1.0), that is, if the first overall quantitative evaluation index exceeds 0.6, it can be determined that the first design scheme and the second design scheme meet the consistency requirement, and the reliability of the first design scheme and the second design scheme is evaluated according to the input, process and output of each link in the first design scheme and the second design scheme; otherwise, if the first overall quantitative evaluation index does not exceed 0.6, it can be determined that the first design scheme and the second design scheme do not meet the consistency requirement, and the preset requirement information is analyzed according to the first pre-training model and the second pre-training model, respectively, and the cycle is repeated until the first overall quantitative evaluation index exceeds the first preset threshold.
[0122] In this embodiment, as the model is gradually upgraded and improved, the first preset threshold can be adjusted accordingly according to the overall evolution of the first pre-training model and the second pre-training model, for example, after multiple upgrades, the first preset threshold can be adjusted to 0.8.
[0123] S330: processing the first design scheme and the second design scheme according to the first credibility evaluation result to obtain a plurality of first positive sample data sets that can be adopted.
[0124] In the embodiment, the first credibility evaluation result is a score of the credibility of the first design scheme and the second design scheme.
[0125] In the embodiment, the data collection and evaluation information service system of the whole life cycle management based on market research can give a reference design result of a reference case based on historical cases and data support of an enterprise big data center, and be used for assisting decision-making, and then the credibility of the scheme of the input, process and output links in each work package is evaluated by manual scoring.
[0126] In the embodiment, the first credibility threshold is set to 7, and if the score in the first credibility evaluation result is greater than or equal to 7, it is determined that the credibility of the first design scheme and the second design scheme meets the requirements; otherwise, if the score in the first credibility evaluation result is less than 7, it is determined that the credibility of the first design scheme and the second design scheme does not meet the requirements.
[0127] In the embodiment, according to the results of the credibility scoring and the first credibility threshold, the results of the input, process and output of each link in the first design scheme and the second design scheme are separated into the first positive sample data that can be adopted and the first negative sample data that cannot be adopted.
[0128] If the first pre-training model and the second pre-training model do not output the design results that can be adopted, the design is improved by manual assistance to form the positive sample data that can be adopted, and all separated sample data are attached with the credibility data determined after manual review.
[0129] Specifically, the first pre-training model and the second pre-training model automatically generate results, but cannot guarantee that the two models together can generate completely satisfactory results, or the credibility of the results generated by the two models is not high, at this time manual assistance is needed, and the result optimized on a better result until the user is satisfied is taken as the final result.
[0130] In the embodiment, the credibility evaluation indexes of the first design scheme and the second design scheme of each WBS work package in the business environment evaluation process research scheme design stage can be processed in the following different cases:
[0131] When the credibility of the result of each link in the WBS work package exceeds the first credibility threshold, the result exceeding the credibility threshold is used as the first positive sample data; when the credibility of the result of each link in the WBS work package exceeds the first credibility threshold, the result with a high credibility score is used as the first positive sample data; when the credibility of the result of each link in the WBS work package does not exceed the first credibility threshold, the result is not used as the positive sample data, at which time the data result needs to be scored again until the first positive sample data is screened out.
[0132] S340: Fuse a plurality of first positive sample data that can be trusted to obtain a first positive sample data set, process the first positive sample data set, obtain a final design scheme, collect market research data according to the final design scheme, and associate the market research data to a newly created market research task.
[0133] In this embodiment, a plurality of first positive sample data that can be trusted is fused into a first positive sample data set, a plurality of first negative sample data that cannot be trusted is fused into a first negative sample data set, and the first positive sample data set with credibility is fused into a final design scheme, so as to facilitate the next operation, and the final design scheme is overlaid to the input, process and output of each link in the first whole-process workflow, for the next field market research.
[0134] In this embodiment, according to the input design scheme, process execution scheme and output result of each sub-process included in the final design scheme, the field research is carried out one by one according to the preset workflow sequence, to collect market research data, and the collected market research data is associated to the input, process and output links of the newly created market research task.
[0135] In the business environment evaluation process, according to the "input-process-output" model definition of each WBS work package, the input is used as the collection boundary, the process defined method and tool are used as the means, and the research activities defined by the process are carried out respectively, and the original data conforming to the output model are collected, which are the market research data in this embodiment.
[0136] In this embodiment, the market research data can include voice, text, video and picture format data, and the collected data in the above formats are associated to the input, process and output links of each sub-process of the market research task, for the generation of the next output content.
[0137] S400: Analyze the market research data according to the first pre-training model and the second pre-training model respectively to obtain first preliminary data results and second preliminary data results.
[0138] In one embodiment of the present application, with reference to Figure 5, step S400 specifically includes the following steps:
[0139] S410: calling the first pre-trained model to process the market research data to obtain a second whole-process workflow, and defining the input, process and output of each link in the second whole-process workflow according to the preset second input module library, second process module library and second output module library.
[0140] Among them, calling the second pre-trained model to analyze and process the market research data can obtain the second whole-process workflow, and the second whole-process workflow in this embodiment is the whole-process workflow of the data result stage, and the input, process and output of each link in the second whole-process workflow can be defined according to the second input module library, the second process module library and the second output module library in the above data collection evaluation information service system, and the creation and process decomposition of the task based on the market research data are completed.
[0141] Among them, facing the business environment evaluation requirements, the minimum work package "input-process-output" model based on work breakdown structure (WBS) is established, and the first pre-trained model is used to process the minimum work package to obtain the parameter or file template requirements of the input, process and output of the minimum work package.
[0142] S420: calling the first pre-trained model to process and analyze the input, process and output of each link in the second whole-process workflow to obtain third input data, third process data and third output data of each link, and fusing the third input data, the third process data and the third output data to obtain the first preliminary data result.
[0143] Among them, in this embodiment, according to the defined input, process and output of each link in the second whole-process workflow, the first pre-trained model is called to process and analyze the final design scheme and the market research data to obtain the required output specified parameter or template requirement, that is, the first preliminary data result.
[0144] Among them, in the "obtaining power" evaluation sub-dimension of the business environment evaluation, respectively for the "input-process-output" model of the six secondary WBS work packages, combining the design scheme of the "input" part and the data collection content of the "process" part, using the information synthesis ability and reasoning ability of the first pre-trained model, the first preliminary data result meeting the "output" requirement is automatically generated.
[0145] S430: calling the second pre-trained model to process and analyze the input, process and output of each link in the second whole-process workflow to obtain fourth input data, fourth process data and fourth output data of each link, and fusing the fourth input data, the fourth process data and the fourth output data to obtain the second preliminary data result.
[0146] In this embodiment, according to the input, process and output of each link in the defined second whole-process workflow, the second pre-trained model is called to process and analyze the final design scheme and market research data to obtain the required output specified parameters or template requirements, that is, the second preliminary data result.
[0147] Among them, in the "power acquisition" evaluation sub-dimension of the business environment evaluation, the "input-process-output" model of the six secondary WBS work packages is combined with the design scheme of the "input" part and the data collection content of the "process" part. The second pre-trained model is used to automatically generate the second preliminary data result that meets the "output" requirement by using the information synthesis ability and reasoning ability of the second pre-trained model.
[0148] S500: According to the second preset rule, the first preliminary data result and the second preliminary data result are analyzed to obtain a second analysis result, and a second positive sample data set of the first preliminary data result and the second preliminary data result is generated according to the second analysis result, and a final data result is obtained according to the second positive sample data set.
[0149] In one embodiment of the present application, with reference to Figure 6 , step S500 specifically includes the following steps:
[0150] S510: Calling the second pre-trained model to perform consistency evaluation on the first preliminary data result and the second preliminary data result to obtain a second consistency evaluation result.
[0151] In this embodiment, for the first preliminary data result obtained in step S420 and the second preliminary data result obtained in step S430, the second pre-trained model is used to perform consistency evaluation on each of the above defined links according to the input, process and output of each link in the defined second whole-process workflow, and finally obtain a second consistency evaluation result. The second consistency evaluation result includes the consistency evaluation result of each link, and then according to the second consistency evaluation result, it is judged whether the first preliminary data result and the second preliminary data result meet the consistency requirement.
[0152] In this embodiment, the basis for consistency evaluation comes from the understanding and reasoning of the second pre-trained model on market research data; and the result of consistency evaluation includes overall qualitative evaluation and pre-set multi-dimension consistency quantitative evaluation. For inconsistent content, the second pre-trained model will give a comparison processing result.
[0153] In the embodiment, the second pre-training model can analyze the market research data and output second preliminary data results, and can also evaluate the consistency of the first preliminary data results and the second preliminary data results by using the semantic understanding and logical reasoning functions of the second pre-training model.
[0154] S520: quantitatively analyze the second consistency evaluation result to obtain a second overall quantitative evaluation index of the first preliminary data results and the second preliminary data results, and if the second overall quantitative evaluation index exceeds a second preset threshold, perform a credibility evaluation on the first preliminary data results and the second preliminary data results to obtain a second credibility evaluation result.
[0155] In the embodiment, the second consistency evaluation result is quantitatively analyzed according to the requirements of the business environment evaluation, and the second quantitative evaluation indices of the input, process and output of each stage in the first preliminary data results and the second preliminary data results can be obtained. Then, the second quantitative evaluation indices of the input, process and output of each stage are weighted and calculated according to the preset weights, and finally the second overall quantitative evaluation index of the first preliminary data results and the second preliminary data results is obtained.
[0156] In the embodiment, the second quantitative evaluation indices of each WBS work package in the business link evaluation task are used to establish a hierarchical dimension weight matrix based on the WBS decomposition hierarchy by using the AHP method, and the second quantitative evaluation indices of each sub-dimension are weighted and calculated to finally obtain the second overall quantitative evaluation index.
[0157] In the embodiment, if the second overall quantitative evaluation index exceeds the second preset threshold, the first preliminary data results and the second preliminary data results are respectively evaluated in terms of the input, process and output of each link to obtain the second credibility evaluation result.
[0158] In the embodiment, the second preset threshold is set to 0.6 (the normalized full score is 1.0), that is, if the second overall quantitative evaluation index exceeds 0.6, it can be determined that the first preliminary data results and the second preliminary data results meet the consistency requirements, and the first preliminary data results and the second preliminary data results are respectively evaluated in terms of the input, process and output of each link. On the contrary, if the second overall quantitative evaluation index does not exceed 0.6, it can be determined that the first preliminary data results and the second preliminary data results do not meet the consistency requirements, and the market research data is collected again according to the final design scheme, and the market research data is analyzed by the first pre-training model and the second pre-training model, respectively, and the cycle is repeated until the second overall quantitative evaluation index exceeds the second preset threshold.
[0159] In this embodiment, as the model is gradually upgraded and improved, the second preset threshold can be adjusted according to the overall evolution of the first pre-training model and the second pre-training model. For example, after several upgrades, the second preset threshold can be adjusted to 0.8.
[0160] S530: Process the first preliminary data result and the second preliminary data result according to the second credibility evaluation result to obtain a plurality of second positive sample data that can be trusted.
[0161] In this embodiment, according to the second credibility evaluation result, the input, process and output results of each link in the first preliminary data result and the second preliminary data result are separated into second positive sample data that can be trusted and second negative sample data that cannot be trusted.
[0162] In this embodiment, according to the existing market research full life cycle management data collection evaluation information service system, reference preliminary data results of reference cases can be given based on historical cases and data support of enterprise big data center, and are used to assist decision-making, and then the credibility of the data results of the input, process and output links in each work package is evaluated by artificial. The evaluation method in this embodiment adopts a scoring method, and the scores are scored according to 0-10, wherein the higher the score, the higher the credibility, 0 represents no credibility, and 10 represents the highest credibility.
[0163] In this embodiment, according to the credibility score and the second credibility threshold, the input, process and output results of each link in the first preliminary data result and the second preliminary data result are separated into second positive sample data that can be trusted and second negative sample data that cannot be trusted.
[0164] If the first pre-training model and the second pre-training model do not output the data results that can be trusted, the data results are improved by artificial assistance to form the second positive sample data that can be trusted, and all separated sample data are attached with the credibility data determined after artificial review.
[0165] In this embodiment, the credibility evaluation index of the first preliminary data result and the second preliminary data result of each WBS work package in the business environment evaluation process data evaluation stage can be processed in the following different cases:
[0166] When the credibility of the data result of each link in the WBS work package exceeds the second credibility threshold, the data result exceeding the second credibility threshold is used as the second positive sample data; when the credibility of the data result of each link in the WBS work package exceeds the second credibility threshold, the data result with a high credibility score is used as the second positive sample data; when the credibility of the data result of each link in the WBS work package does not exceed the second credibility threshold, the data result is not used as the second positive sample data, at which time the data result needs to be scored again until the second positive sample data is screened out.
[0167] In this embodiment, the credibility threshold is set to 7. Of course, according to the actual use, the set value of the credibility threshold can also be adjusted accordingly, and the present application does not limit this.
[0168] S540: Fuse a plurality of second positive sample data that can be trusted to obtain a second positive sample data set, and process the second positive sample data set to obtain a final data result.
[0169] In this embodiment, the second positive sample data that can be trusted is fused into a second positive sample data set, the second negative sample data that cannot be trusted is fused into a second negative sample data set, and the second positive sample data set with credibility is fused into a final data result, so as to facilitate the next operation.
[0170] S600: Determine the new data content in the blockchain according to the first positive sample data set and the second positive sample data set, generate a new block according to the new data content, and respond to and authenticate the new block.
[0171] In one embodiment of the present application, with reference to Figure 7 and Figure 9 Step S600 specifically includes the following steps:
[0172] S610: The distributed node obtains authorization from the CA root node server.
[0173] In this embodiment, the distributed node in the blockchain is a user-side intelligent terminal authenticated by a digital CA certificate; the intelligent terminal in this embodiment includes at least one of a computer, a tablet, a mobile phone, a workstation and a server; and each distributed node stores a key pair of the node, including a public key pair and a private key pair, and obtains authorization and manages the key pair through a digital CA certificate.
[0174] In this embodiment, the authorized distributed node in the enterprise network domain obtains a digital certificate authorization from the enterprise CA root node server to complete the identity authentication of the device.
[0175] S620: If the distributed node obtains the first positive sample data set and the second positive sample data set with credibility, a new block is automatically generated by the distributed node, the new block is encrypted by the key pair stored by the distributed node, and broadcasted to other distributed nodes.
[0176] In this embodiment, if the distributed node can obtain the first positive sample data set and the second positive sample data set with credibility, a new block in the blockchain is automatically generated by the distributed node, and the new block is encrypted by the key pair stored by the distributed node, and then broadcasted to other distributed nodes in the enterprise management domain, so that other distributed nodes can obtain the data content in the new block.
[0177] In this embodiment, the other distributed nodes in the enterprise management domain include other distributed nodes of the public chain in the blockchain and other distributed nodes of the private chain in the blockchain.
[0178] In one embodiment of the present application, reference is made to Figure 8 and Figure 9 The step S600 further includes the following steps:
[0179] S630: The other distributed nodes of the private chain of the blockchain authenticate the new block according to the consensus mechanism of the private chain. After the authentication is passed, the new block takes effect, and the first positive sample data set and the second positive sample data set are automatically linked to the tail of the private chain.
[0180] According to the consensus mechanism of the private chain, the other distributed nodes of the private chain of the blockchain authenticate the new block, and after the authentication is passed, the new block takes effect, and then the obtained first positive sample data set and the second positive sample data set are automatically linked to the tail of the private chain, and finally the positive sample data is completed. The chain process.
[0181] S640: After the new block takes effect, the CA root node server re-packages the data content of the new block, encrypts the new block by the key pair stored by the CA root node server, and broadcasts to the other distributed nodes of the public chain of the blockchain.
[0182] After the new block takes effect, the CA root node server re-packages the data content of the new block, and encrypts the new block according to the key pair stored by the CA root node server, and then broadcasts to the other distributed nodes of the public chain of the blockchain, so that the other distributed nodes of the public chain can obtain the data content of the new block.
[0183] In this embodiment, the re-packaging of the content of the new block is the linking of the data content to the tail of the public chain mentioned below.
[0184] S650: Other distributed nodes of the public chain authenticate the new block according to the consensus mechanism of the public chain. After the authentication is passed, the new block takes effect, and the first positive sample data set and the second positive sample data set are automatically linked to the tail of the public chain.
[0185] According to the consensus mechanism of the public chain, other distributed nodes of the public chain of the block chain authenticate the new block, and the new block takes effect after the authentication is passed. Then the obtained first positive sample data set and the second positive sample data set are automatically linked to the tail of the public chain, and finally the positive sample data is completed.
[0186] The public chain provides a root node digital certificate authority for the CA root node server of the public chain, generates a root CA certificate, and completes the identity authentication and trust of the enterprise CA root node server.
[0187] In this embodiment, the data content contained in the new block in the chain process of the positive sample data of the research design stage and the data evaluation stage can include:
[0188] The result matrix of each WBS work package "input" link after credibility evaluation; The result matrix of each WBS work package "process" link after credibility evaluation; The result matrix of each WBS work package "output" link after credibility evaluation; The credibility evaluation matrix of each WBS work package "input" link data; The credibility evaluation matrix of each WBS work package "process" link data; The credibility evaluation matrix of each WBS work package "output" link data; The model input associated with each WBS work package "input" link; The model input associated with each WBS work package "process" link; The model input associated with each WBS work package "output" link.
[0189] In one embodiment of the present application, the first negative sample data set of the first design scheme and the second design scheme can also be generated according to the first analysis result, the second negative sample data set of the first preliminary data result and the second preliminary data result can be generated according to the second analysis result, and the data content of the new block is respectively fused into the first positive sample data set and the second positive sample data set. The first positive sample data set, the second positive sample data set, the first negative sample data set and the second negative sample data set are fused into the market research special sample data set.
[0190] Wherein, the first negative sample data set can be obtained by fusing the first negative sample data, and the second negative sample data set can be obtained by fusing the second negative sample data; and the data content of the new block is fused into the first positive sample data set with credibility and the second positive sample data set with credibility respectively, and the first positive sample data set and the second positive sample data set are fused into the market research special sample data set, at the same time, the first negative sample data set and the second negative sample data set are fused into the market research special sample data set, and the market research special sample data set and the open data set are fused again to obtain a new initial training data set, and the new initial training data set is used to train the models with different architectures respectively, so that the new first pre-training model and the second pre-training model can be obtained to improve the model training performance of the first pre-training model and the second pre-training model.
[0191] The implementation principle of the embodiments of the present application is that the first pre-training model and the second pre-training model can be obtained by training the models with different architectures according to the obtained initial training data set, and the two pre-training models are called to analyze the preset demand information respectively to obtain two design schemes, the second pre-training model is called to analyze the two design schemes to obtain a first analysis result, and the first positive sample data set is generated according to the first analysis result, and then the final design scheme is generated according to the first positive sample data set, and the market research data is collected according to the final design scheme; the two pre-training models are called to analyze the market research data respectively, and two preliminary data results can be obtained, the second pre-training model is called to analyze the two preliminary data results to obtain a second analysis result, and the second positive sample data set is generated according to the second analysis result, and then the final data result is generated according to the second positive sample data set, and then the new block is generated according to the first positive sample data set and the second positive sample data set, and the new block is responded and authenticated to complete the data content in the new block to be chained to the block chain; so that the data processing efficiency and data accuracy in the data conclusion stage can be improved.
[0192] The embodiment of the application discloses a blockchain data acquisition device based on a multi-architecture NLP pre-training model, which specifically comprises a model construction module, a first generation module, a first analysis module, a second generation module, a second analysis module and a third generation module; wherein the model construction module is used for obtaining initial training data sets in response to a request, and training models of different architectures according to the initial training data sets to obtain first and second pre-training models; the first generation module is used for training preset demand information according to the first and second pre-training models to obtain first and second design schemes; the first analysis module is used for analyzing the first and second design schemes according to a first preset rule to obtain a first analysis result, generating a first positive sample data set of the first and second design schemes according to the first analysis result, obtaining a final design scheme according to the first positive sample data set, and collecting market research data according to the final design scheme; the second generation module is used for analyzing market research data according to the first and second pre-training models to obtain first and second preliminary data results; the second analysis module is used for analyzing the first and second preliminary data results according to a second preset rule to obtain a second analysis result, generating a second positive sample data set of the first and second preliminary data results according to the second analysis result, and obtaining a final data result according to the second positive sample data set; and the third generation module is used for determining new data content in the blockchain according to the first and second positive sample data sets, generating a new block according to the new data content, and responding to and authenticating the new block.
[0193] In one embodiment of the application, the device adopts the blockchain data acquisition method based on the multi-architecture NLP pre-training model of the above-mentioned embodiment in specific application. The specific application content of the device is the same as the method content, and thus will not be described here.
[0194] The embodiment of the application discloses a terminal, which comprises a memory, a processor and a computer program stored in the memory and capable of running on the processor, and the processor adopts the blockchain data acquisition method based on the multi-architecture NLP pre-training model of the above-mentioned embodiment when executing the computer program.
[0195] The embodiment of the application discloses a computer readable storage medium, and the computer readable storage medium stores a computer program, wherein the computer program is executed by a processor, and the blockchain data acquisition method based on the multi-architecture NLP pre-training model of the above-mentioned embodiment is adopted.
[0196] The embodiments of the present application are preferred embodiments of the present application, and are not intended to limit the protection scope of the present application, wherein the same parts are denoted by the same reference numerals. Therefore, any equivalent changes made according to the structure, shape and principle of the present application should be covered within the protection scope of the present application.
Claims
1. A blockchain data collection method based on a multi-architecture NLP pre-training model, characterized in that, The method comprises the following steps: in response to a request, obtaining an initial training data set, and training different architecture models according to the initial training data set to obtain a first pre-training model and a second pre-training model; training the first pre-training model and the second pre-training model according to preset demand information to obtain a first design scheme and a second design scheme; analyzing the first design scheme and the second design scheme according to a first preset rule to obtain a first analysis result, generating a first positive sample data set of the first design scheme and the second design scheme according to the first analysis result, obtaining a final design scheme according to the first positive sample data set, and collecting market research data according to the final design scheme; analyzing the market research data according to the first pre-training model and the second pre-training model to obtain a first preliminary data result and a second preliminary data result; analyzing the first preliminary data result and the second preliminary data result according to a second preset rule to obtain a second analysis result, generating a second positive sample data set of the first preliminary data result and the second preliminary data result according to the second analysis result, and obtaining a final data result according to the second positive sample data set; determining new data content in the blockchain according to the first positive sample data set and the second positive sample data set, generating a new block according to the new data content, and responding to and authenticating the new block; the method for obtaining an initial training data set in response to a request, and training different architecture models according to the initial training data set to obtain a first pre-training model and a second pre-training model, specifically comprising: obtaining an open data set, generating a market research special sample data set according to market research related data, and fusing the open data set and the market research special sample data set into the initial training data set; calling the initial training data set to train a GPT model of a Decoder structure of a Transformer to obtain the first pre-training model, and testing the first pre-training model; calling the initial training data set to train a BERT model of an Encoder structure of a Transformer to obtain the second pre-training model, and testing the second pre-training model; the method for training the first pre-training model and the second pre-training model according to preset demand information to obtain a first design scheme and a second design scheme, specifically comprising: calling the first pre-training model to process the preset demand information to obtain a first whole-process workflow, and defining the input, process and output of each link in the first whole-process workflow according to a preset first input module library, a first process module library and a first output module library; The first pre-training model is called to process and analyze the input, process and output of each link in the first whole-process workflow, to obtain first input data, first process data and first output data of each link, and to fuse the first input data, the first process data and the first output data, to obtain the first design scheme; The second pre-training model is called to process and analyze the input, process and output of each link in the first whole-process workflow, to obtain second input data, second process data and second output data of each link, and to fuse the second input data, the second process data and the second output data, to obtain the second design scheme; The first positive sample data set and the second positive sample data set are determined according to the first positive sample data set and the second positive sample data set, and a new block is generated according to the new data content, specifically including: The distributed node obtains authorization from the CA root node server; If the distributed node obtains the first positive sample data set and the second positive sample data set with credibility, a new block is automatically generated through the distributed node, the new block is encrypted by using the key stored in the distributed node, and the new block is broadcasted to other distributed nodes; The new block is responded and authenticated, specifically including: The other distributed nodes of the private chain of the blockchain authenticate the new block according to the consensus mechanism of the private chain, and after the authentication is passed, the new block takes effect, and the first positive sample data set and the second positive sample data set are automatically linked to the tail of the private chain; When the new block takes effect, the CA root node server re-packages the data content of the new block, encrypts the new block by using the password stored in the CA root node server, and broadcasts the new block to other distributed nodes of the public chain of the blockchain; The other distributed nodes of the public chain authenticate the new block according to the consensus mechanism of the public chain, and after the authentication is passed, the new block takes effect, and the first positive sample data set and the second positive sample data set are automatically linked to the tail of the public chain.
2. The method of claim 1, wherein the method is based on a multi-architecture NLP pre-training model. The first design scheme and the second design scheme are analyzed according to the first preset rule to obtain a first analysis result, a first positive sample data set of the first design scheme and the second design scheme is generated according to the first analysis result, a final design scheme is obtained according to the first positive sample data set, market research data is collected according to the final design scheme, specifically including: The second pre-training model is called to evaluate the consistency of the first design scheme and the second design scheme, to obtain a first consistency evaluation result; The first consistency evaluation result is quantitatively analyzed to obtain a first overall quantitative evaluation index of the first design scheme and the second design scheme, and if the first overall quantitative evaluation index exceeds a first preset threshold, a credibility evaluation of the first design scheme and the second design scheme is performed to obtain a first credibility evaluation result; According to the first credibility evaluation result, the first design scheme and the second design scheme are processed to obtain a plurality of first positive sample data that can be adopted; The first positive sample data set is obtained by fusing a plurality of the first positive sample data that can be adopted, and the first positive sample data set is processed to obtain the final design scheme. Market research data is collected according to the final design scheme, and the market research data is associated with a new market research task.
3. The method of claim 2, wherein the method is based on a multi-architecture NLP pre-training model. According to the first pre-training model and the second pre-training model, the market research data is analyzed to obtain first preliminary data results and second preliminary data results, specifically including: The first pre-training model is called to process the market research data to obtain a second whole-process workflow, and a second input module library, a second process module library and a second output module library are defined according to a preset to define the input, process and output of each link in the second whole-process workflow; The first pre-training model is called to process and analyze the input, process and output of each link in the second whole-process workflow to obtain third input data, third process data and third output data of each link, and the third input data, the third process data and the third output data are fused to obtain the first preliminary data results; The second pre-training model is called to process and analyze the input, process and output of each link in the second whole-process workflow to obtain fourth input data, fourth process data and fourth output data, and the fourth input data, the fourth process data and the fourth output data are fused to obtain the second preliminary data results.
4. The method of claim 3, wherein the method is based on a multi-architecture NLP pre-training model. According to the second pre-training model, the first preliminary data results and the second preliminary data results are analyzed to obtain a second analysis result, and a second positive sample data set of the first preliminary data results and the second preliminary data results is generated according to the second analysis result, and a final data result is obtained according to the second positive sample data set, specifically including: The second pre-training model is called to evaluate the consistency of the first preliminary data results and the second preliminary data results to obtain a second consistency evaluation result; The second consistency evaluation result is quantitatively analyzed to obtain a second overall quantitative evaluation index of the first preliminary data results and the second preliminary data results, and if the second overall quantitative evaluation index exceeds a second preset threshold, the first preliminary data results and the second preliminary data results are evaluated for credibility to obtain a second credibility evaluation result; According to the second credibility evaluation result, the first preliminary data results and the second preliminary data results are processed to obtain a plurality of second positive sample data that can be adopted; The second positive sample data set is obtained by fusing a plurality of the second positive sample data that can be adopted, and the second positive sample data set is processed to obtain the final data result.
5. The method of claim 1, wherein the method is based on a multi-architecture NLP pre-training model. According to the first analysis result, a first negative sample data set of the first design scheme and the second design scheme is generated, according to the second analysis result, a second negative sample data set of the first preliminary data result and the second preliminary data result is generated, and the data content of the new block is fused into the first positive sample data set and the second positive sample data set respectively, and the first positive sample data set, the second positive sample data set, the first negative sample data set and the second negative sample data set are fused into the market research special sample data set.
6. A blockchain data collection device based on a multi-architecture NLP pre-training model, characterized in that, The device comprises a blockchain data acquisition method based on a multi-architecture NLP pre-training model according to any one of claims 1-5, and the device comprises: A model construction module is configured to obtain an initial training data set in response to a request, and train models of different architectures according to the initial training data set to obtain a first pre-training model and a second pre-training model. A first generation module is configured to train preset demand information according to the first pre-training model and the second pre-training model to obtain a first design scheme and a second design scheme. A first analysis module is configured to analyze the first design scheme and the second design scheme according to a first preset rule to obtain a first analysis result, generate a first positive sample data set of the first design scheme and the second design scheme according to the first analysis result, obtain a final design scheme according to the first positive sample data set, and acquire market research data according to the final design scheme. A second generation module is configured to analyze the market research data according to the first pre-training model and the second pre-training model to obtain a first preliminary data result and a second preliminary data result. A second analysis module is configured to analyze the first preliminary data result and the second preliminary data result according to a second preset rule to obtain a second analysis result, generate a second positive sample data set of the first preliminary data result and the second preliminary data result according to the second analysis result, and obtain a final data result according to the second positive sample data set. A third generation module is configured to determine new data content in a blockchain according to the first positive sample data set and the second positive sample data set, generate a new block according to the new data content, and respond to and authenticate the new block.
Citation Information
Patent Citations
Intelligent generation method and system of planning scheme based on machine learning
CN109472390A
Form generation method and device, computer equipment and storage medium
CN113052262A
Natural language processing method and device based on multiple tools and medium
CN114580387A
Commercial vehicle part design automatic evaluation method, device and equipment and storage medium
CN115907298A