A method, device and equipment for allocating resources of a large model based on a blockchain
By employing methods of data cleaning and multi-index evaluation, combined with smart contracts and blockchain technology, the problem of inconsistent data quality in the training of large language models was solved, enabling precise allocation of resources and fair incentives, thereby improving training effectiveness and resource utilization efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 中国移动通信集团云南有限公司
- Filing Date
- 2026-03-02
- Publication Date
- 2026-07-03
AI Technical Summary
During the training of large language models, training data comes from a wide range of sources but varies in quality. There is a large amount of inefficient, invalid, or even biased data, which leads to poor training results and wasted computing resources.
By receiving and cleaning the raw data sent by the data provider, generating data summary information, and using multi-indicator evaluation and smart contracts, the data is evaluated for its gain, rarity, and performance metrics. Resources are allocated to the data provider only when the model performance improvement reaches a preset threshold, and the relevant information is stored on the blockchain.
It achieves accurate and fair allocation of large model resources, improves resource utilization efficiency and incentive mechanisms, ensures the quality and impartiality of training data, and reduces illusions and biases.
Smart Images

Figure CN122332854A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of blockchain technology, and in particular to a method, apparatus, and device for allocating large-scale blockchain-based resources. Background Technology
[0002] Large language models are typically pre-trained and fine-tuned based on massive amounts of data, and their capabilities largely depend on the scale, coverage, and quality of the training data.
[0003] Currently, in the training process of large language models, specific activation functions or reward signals are designed to guide the model to learn and optimize its output behavior, making it more in line with human expectations or specific task goals. However, in the training process of large language models, the sources of training data are wide but of varying quality. A large amount of inefficient, invalid, or even biased data is introduced into the training process, resulting in poor training effects, the potential for illusions or biases, and a waste of computing resources. Summary of the Invention
[0004] This disclosure provides a blockchain-based method, apparatus, and device for allocating large-scale model resources, which, through multi-indicator evaluation and smart contracts, achieves accurate and fair allocation of large-scale model resources, thereby improving resource utilization and incentive effects.
[0005] In a first aspect, embodiments of this disclosure provide a method for allocating large-scale blockchain-based resources, the method comprising: Receive first material data sent by at least one data provider, clean the first material data to obtain second material data, and determine the data summary information of the second material data, wherein the first material data includes at least text data, audio data and / or video data, and the second material data includes at least a portion of the data in the first material data; The data summary information is input into the target model in batches, and the indicator data of each data summary information under multiple evaluation indicators are determined. The evaluation indicators include gain evaluation indicators, rarity evaluation indicators and performance indicators. The indicator data includes information gain corresponding to the gain evaluation indicator, data rarity attribute corresponding to the rarity evaluation indicator and model performance improvement attribute corresponding to the performance indicator. For multiple data summary information, the indicator data of the data summary information is sent to the smart contract so that when the performance improvement attribute of the model reaches a preset attribute threshold, the data provider corresponding to the data summary information is allocated a first resource. The data digest information, the data provider, the first resource, and the sending address and time information of the data digest information are stored on the blockchain.
[0006] Secondly, embodiments of the present invention also provide a blockchain-based large-scale resource allocation device, the device comprising: A data summary information determination module is used to receive first material data sent by at least one data provider, clean and process the first material data to obtain second material data, and determine the data summary information of the second material data, wherein the first material data includes at least text data, audio data and / or video data, and the second material data includes at least a portion of the data in the first material data; The indicator data determination module is used to input the data summary information into the target model in batches and determine the indicator data of each data summary information under multiple evaluation indicators. The evaluation indicators include gain evaluation indicators, rarity evaluation indicators and performance indicators. The indicator data includes information gain corresponding to the gain evaluation indicator, data rarity attribute corresponding to the rarity evaluation indicator and model performance improvement attribute corresponding to the performance indicator. The first resource allocation module is used to send the indicator data of the data summary information to the smart contract for multiple data summary information, so as to allocate the first resource to the data provider corresponding to the data summary information when it is determined that the performance improvement attribute of the model reaches a preset attribute threshold. The on-chain storage module is used to store the data digest information, the data provider, the first resource, and the sending address and time information of the data digest information on the blockchain.
[0007] Thirdly, embodiments of the present invention also provide an electronic device, the electronic device comprising: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the blockchain-based large model resource allocation method as described in any embodiment of the present invention.
[0008] Fourthly, embodiments of the present invention also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the blockchain-based large-scale resource allocation method as described in any of the embodiments of the present invention.
[0009] Fifthly, embodiments of the present invention also provide a computer program product, including a computer program, characterized in that, when executed by a processor, the computer program implements the blockchain-based large model resource allocation method as described in any embodiment of the present invention.
[0010] The technical solution of this disclosure firstly receives first source data sent by at least one data provider, cleans and processes the first source data to obtain second source data, and determines the data summary information of the second source data. Then, the data summary information is batch-input into the target model, and the indicator data for each data summary information under multiple evaluation metrics are determined. Further, for multiple data summary information, the indicator data of the data summary information is sent to a smart contract so that when the model performance improvement attribute reaches a preset attribute threshold, a first resource is allocated to the data provider corresponding to the data summary information. Finally, the data summary information, data provider, first resource, and the sending address and time information of the data summary information are stored on the blockchain. This solves the problems in the prior art where, during the training process of large language models, training data sources are wide-ranging but of varying quality, with a large amount of inefficient, invalid, or even biased data being introduced into the training process, resulting in poor training effects, potential illusions or biases, and wasted computing resources. This disclosure embodiment achieves accurate and fair allocation of large model resources through multi-metric evaluation and smart contracts, thereby improving resource utilization and incentive effects. Attached Figure Description
[0011] To more clearly illustrate the technical solutions of exemplary embodiments of the present invention, the accompanying drawings used in describing the embodiments are briefly introduced below. Obviously, the accompanying drawings described are only a portion of the drawings of the embodiments to be described in this invention, and not all of the drawings. For those skilled in the art, other drawings can be obtained from these drawings without any creative effort.
[0012] Figure 1 This is a flowchart illustrating a large-scale resource allocation method based on blockchain provided in this embodiment of the disclosure; Figure 2 This is a flowchart illustrating a large-scale resource allocation method based on blockchain provided in this embodiment of the disclosure; Figure 3 This is a schematic diagram of a large-scale blockchain-based resource allocation method provided in an embodiment of this disclosure; Figure 4 A schematic diagram of a blockchain-based large-scale resource allocation device provided in an embodiment of this disclosure; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0013] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention, and not all of the structures.
[0014] Before introducing the technical solutions provided in the embodiments of this disclosure, the application scenarios can be illustrated first. The technical solutions provided in the embodiments of this disclosure can be applied to scenarios involving resource allocation for large models. Based on the technical solutions in the embodiments of this disclosure, the accurate and fair allocation of resources for large models is achieved through multi-index evaluation and smart contracts, thereby improving resource utilization and incentive effects.
[0015] Example 1 Figure 1 This is a flowchart illustrating a blockchain-based large model resource allocation method provided in this embodiment. This embodiment is applicable to the scenario of allocating resources for large models. The method can be executed by a blockchain-based large model resource allocation device, which can be implemented in the form of software and / or hardware. The hardware can be a mobile electronic device, which can execute the blockchain-based large model resource allocation method provided in this technical solution.
[0016] like Figure 1 As shown, the method includes: S110. Receive first material data sent by at least one data provider, clean the first material data to obtain second material data, and determine the data summary information of the second material data.
[0017] The first material data includes at least text data, audio data, and / or video data, and the second material data includes at least a portion of the data in the first material data.
[0018] It should be noted that a data provider refers to the entity that uploads or submits the raw data used for training the target model. A data provider can be an individual user, institution, data platform, or any entity capable of submitting data and associating its identity through its account or public key address. The first source data refers to the collection of raw source data sent by the data provider. The first source data can be a single or multiple data records. The second source data refers to a standardized dataset obtained after preprocessing the first source data, such as cleaning, and can be used for subsequent training. Furthermore, the second source data includes at least a portion of the data from the first source data. For example, after removing duplicate, low-quality, and non-compliant fragments from the first source data, the remaining valid portion constitutes the second source data. Data digest information is used to characterize the key features or unique identifier of the second source data without exposing its plaintext content. Data digest information may include: metadata such as hash fingerprints, data type, data length, timestamps, and source addresses; feature vectors; and other compressed representations used for value assessment. It should also be noted that text data refers to data that expresses semantic content in the form of sequences of characters, words, or symbols. Text data can include natural language text, code, structured text fields, or combinations thereof. Text data can originate from media such as dialogues, documents, web pages, and comments. Audio data refers to data recorded or represented in the form of sound signals. Audio data can include speech or ambient sound. Audio data can be represented using waveforms, sequences of sample points, or encoded and compressed audio file formats. Video data refers to data composed of a continuous sequence of image frames. Video data can be represented as a frame sequence, an encoded video file, or segments thereof. Video data is used to present dynamic visual content.
[0019] Specifically, the system receives text, audio, and / or video data from at least one data provider. It performs cleaning, noise reduction, and deduplication on the first data source to avoid lowering training effectiveness and wasting computational resources, resulting in second data source that is more suitable for model training. It generates data summary information, such as hashes or metadata, for the second data source to uniquely identify it and link it to the provider's source, facilitating subsequent verification.
[0020] Optionally, if the first material data is text data and / or audio data, data deduplication is performed based on the word vectors of the first material data to obtain the second material data of text data; and / or, if the first material data is video data, the first feature information of the first material data is determined based on the feature extraction model, and based on the first feature information and the historical feature information of the historical material data, it is determined whether to use the first material data as the second material data.
[0021] In this context, word vectors refer to a set of numerical vectors that map the text content corresponding to the first source data. Word vectors represent the semantic features of the first source data. Data deduplication involves identifying and deleting duplicate or highly similar data records, retaining only one copy or the copy with higher quality. A feature extraction model is a model used to extract numerical features from video data that can be used for comparison and recognition. For example, based on a feature extraction model, keyframe visual features or temporal features of the first source data can be extracted. First feature information is a feature representation calculated from the video data. For example, first feature information can be a vector, a feature sequence, or a feature signature.
[0022] It should be noted that historical footage data refers to a collection of footage data that has been received, processed, and retained. Historical feature information is the feature representation corresponding to historical footage data, used for similarity comparison with the first feature information of the video data.
[0023] Specifically, when the first source data is text and / or audio data, corresponding word vectors are first generated for this data. The semantic similarity between different data is measured based on the word vectors, allowing for the determination of which data are repetitive or highly similar. After deleting or filtering duplicate data, the remaining content constitutes the second source data, consisting of text and / or audio data. When the first source data is video data, the first feature information of the video data is first calculated using a feature extraction model. Then, this first feature information is compared with the historical feature information of historical source data. If the comparison result shows that the video data is repetitive or highly similar to historical source data, the video data is not included in the second source data; if the comparison result shows that the video data has sufficient difference or novelty, the video data is retained as the second source data.
[0024] S120. Input the data summary information into the target model in batches and determine the indicator data of each data summary information under multiple evaluation indicators.
[0025] The evaluation metrics include gain evaluation metrics, rarity evaluation metrics, and performance metrics. The metric data includes information gain corresponding to the gain evaluation metrics, data rarity attributes corresponding to the rarity evaluation metrics, and model performance improvement attributes corresponding to the performance metrics.
[0026] It's important to note that the target model is the large model ontology used for training or evaluation, used to determine the training value of the data. Evaluation metrics are dimensions or calculation standards used to measure the value of a piece of data and the extent of that value. Metric data are the specific numerical results calculated under each evaluation metric. Gain metrics measure how much new information or uncertainty a piece of data brings. Rarity metrics measure whether a piece of data is uncommon and novel within the existing dataset. Performance metrics measure the actual improvement that a piece of data brings to the target model's performance.
[0027] It's also worth noting that information gain refers to how much the uncertainty of the target model's output is reduced after introducing the data. A higher information gain generally indicates more useful data. Data scarcity refers to the frequency of the data's occurrence in the training set. The rarer the data, the higher the corresponding data scarcity value. Model performance improvement refers to how much the target model's metrics, such as accuracy, improve on the validation or test set after training with the data, compared to before training.
[0028] Specifically, data summaries from multiple datasets are fed into the target model in batches for calculation. For each data summary, values corresponding to multiple dimensions such as gain, rarity, and performance improvement are calculated. These values are then used as the performance metrics for that data point under each evaluation indicator.
[0029] Optionally, the data summary information is input into the target model, and the probability information of the data summary information is determined based on the first function in the target model; the information gain is determined based on the probability information; and the data rarity attribute is determined based on the frequency of occurrence of the data summary information in the training set.
[0030] Here, the first function is the computational function within the target model used to map the input to the output probability distribution. For example, the first function can be a Softmax output layer or a probability prediction head. Probability information is the probability distribution output by the target model for a given piece of data summary information. For example, probability information can be the probability of each category or the probability of each word / unit. Frequency of occurrence refers to the number of times or proportion of the data summary information or its equivalent or similar summary features appear in the training dataset.
[0031] It should be noted that the first function can be the mutual information formula. The information gain can be determined based on the mutual information formula. The mutual information formula is: ; in, The entropy output by the target model can measure the uncertainty of random variables. This indicates that no data summary information is used. At that time, the model output The inherent uncertainty.
[0032] ; in, The target model output Take a certain value The probability of. The larger the value, the more chaotic the output of the target model and the higher the uncertainty. Indicates that the known data summary information Then, the target model outputs... The remaining uncertainty. The calculation formula is: ; in, It is known data summary information Take a certain value The probability of. It is known hour, Pick The conditional probability. The smaller the value, the more information it contains in the data summary. It can significantly reduce The uncertainty. Information gain can be used to... This indicates that the information gain is and The difference represents the data summary information. Output for target model The resulting increase in certainty. The larger the data summary information right The stronger the predictive ability, the greater the core contribution of the data to the training of the target model.
[0033] It's also important to note that when determining data rarity attributes, we count the frequency or proportion of the data summary information appearing in the training set. Then, the frequency of the data summary information in the training set can be mapped to the data rarity attribute. The fewer times the data summary information appears in the training set, the higher the data rarity attribute. For example, an inverse proportion or a monotonically decreasing function can be used to calculate the data rarity attribute.
[0034] In this embodiment, the previously determined historical performance attributes are obtained, and the data summary information is used as training data to train the target model. The test set is then processed based on the target model to obtain the current performance attributes. Based on the current performance attributes and the historical performance attributes, the performance improvement attribute of the model is determined.
[0035] Among them, the historical performance attribute is determined by processing the test set based on the target model obtained from the previous round of training.
[0036] It should be noted that historical performance attributes refer to the performance metrics obtained by evaluating the target model using the same test set after the previous training round. For example, historical performance attributes can be accuracy, F1 score, or loss value. Current performance attributes refer to the performance metrics obtained by evaluating the target model using the newly added training data in the current round using the same test set.
[0037] Specifically, the historical performance attributes recorded in the previous round are retrieved. Data corresponding to the current round's data summary information is added to the training process as training data to update the target model. After training, the updated target model is used for inference or evaluation on the test set, and the results obtained from the test set evaluation are used as the current performance attribute. The current performance attribute is compared with the historical performance attribute, and the difference or relative change between the two is calculated as the model performance improvement attribute.
[0038] S130. For multiple data summary information, send the indicator data of the data summary information to the smart contract so that when the performance improvement attribute of the model reaches the preset attribute threshold, the first resource is allocated to the data provider corresponding to the data summary information.
[0039] Smart contracts are automated programs deployed on the blockchain, used to verify, record transactions, and distribute rewards according to preset rules, publicly recording the results in an immutable manner. Preset attribute thresholds are pre-defined minimum performance improvement standards used to determine whether the current data has brought sufficient performance improvement to the target model to trigger rewards. Primary resources are the rewards or resource rights allocated to the data provider. Primary resources can be on-chain points / tokens, computing power quotas, API call limits, service usage rights, or other quantifiable resources.
[0040] Specifically, for the metric data corresponding to multiple data summary messages, these metric data are packaged and sent to the smart contract. The smart contract checks whether the model performance improvement attribute reaches a preset attribute threshold. If the model performance improvement attribute reaches the preset attribute threshold, the smart contract allocates the first resource to the data provider corresponding to the data summary message; if the model performance improvement attribute does not reach the preset attribute threshold, the smart contract does not trigger or reduces the allocation of the first resource.
[0041] S140. Store the data summary information, data provider, first resource, and the sending address and time information of the data summary information on the blockchain.
[0042] The sending address refers to the on-chain account address or network identity identifier that submitted the data digest information, used to identify who initiated the submission. The time information refers to the timestamp when the data digest information was submitted to the blockchain.
[0043] Specifically, the data summary information and its corresponding data provider identity are recorded on the blockchain. The quantity or type of the first resource to be allocated is also recorded on the blockchain. Simultaneously, the sending address and the time of submission of the data summary information are recorded. During on-chain storage, the data summary information is organized into a fixed-format summary value and its identifier field. The data provider identifier is determined as on-chain identity information, and the first resource is organized into a recordable parameter. The sending address is obtained and written into the field as the account address initiating this on-chain transaction. The time information is obtained and written into the field using the blockchain transaction timestamp. Then, the above fields are combined into structured data for an on-chain record. The notarization or accounting method of the smart contract is called, and the structured data is written as a parameter into the contract state or event log. The transaction is signed using the private key corresponding to the sending address to prove the authenticity and validity of the submission source. The signed transaction is broadcast to the blockchain network and awaits node packaging and consensus confirmation. After transaction confirmation, this information is stored on the chain in an immutable and traceable form and can be retrieved through the transaction hash or contract query interface.
[0044] The technical solution of this disclosure firstly receives first source data sent by at least one data provider, cleans and processes the first source data to obtain second source data, and determines the data summary information of the second source data. Then, the data summary information is batch-input into the target model, and the indicator data for each data summary information under multiple evaluation metrics are determined. Further, for multiple data summary information, the indicator data of the data summary information is sent to a smart contract so that when the model performance improvement attribute reaches a preset attribute threshold, a first resource is allocated to the data provider corresponding to the data summary information. Finally, the data summary information, data provider, first resource, and the sending address and time information of the data summary information are stored on the blockchain. This solves the problems in the prior art where, during the training process of large language models, training data sources are wide-ranging but of varying quality, with a large amount of inefficient, invalid, or even biased data being introduced into the training process, resulting in poor training effects, potential illusions or biases, and wasted computing resources. This disclosure embodiment achieves accurate and fair allocation of large model resources through multi-metric evaluation and smart contracts, thereby improving resource utilization and incentive effects.
[0045] Example 2 Figure 2 This is a flowchart illustrating the blockchain-based large-scale model resource allocation method provided in this embodiment of the invention. Based on the aforementioned embodiments, it provides a detailed explanation of adjusting the weights of multiple evaluation indicators. For specific implementation details, please refer to the technical solution of this embodiment. Technical terms that are the same as or corresponding to those in the above embodiments will not be repeated here.
[0046] like Figure 2 As shown, the method specifically includes the following steps: S210. Adjust the weights of multiple evaluation indicators based on the model performance improvement attribute in the indicator data.
[0047] Among them, the indicator weight refers to the relative importance coefficient of each evaluation indicator in the comprehensive scoring or decision-making process, which is used to determine the magnitude of the impact of each indicator on the final result.
[0048] Specifically, based on the model performance improvement attribute in the indicator data, the weighting of gain evaluation indicators, rarity evaluation indicators, and performance indicators in the overall score is reallocated. By adjusting the weighting of each evaluation indicator in the overall score, the overall score is aligned with the actual improvement in model performance by the data, thus more accurately reflecting the true contribution of the data. After the weighting adjustment, resources are more inclined to be allocated to data providers that truly bring about performance improvements, thereby improving the overall return on investment in training.
[0049] S220. For data summary information of the same category, determine the cumulative contribution attribute of the same category based on the information gain, data rarity attribute, and model performance improvement attribute of the data summary information, and adjust the data weight of the same category based on the cumulative contribution attribute.
[0050] In this context, data summary information of the same category refers to a group of data summary information with the same label, task type, or business domain affiliation. For example, data summary information of the same category could be data summary information belonging to the same fault type, the same intent category, or the same topic domain. The cumulative contribution attribute of a certain type refers to the total score of the cumulative contribution of all data summary information under that category to the improvement of the target model. The cumulative contribution attribute is usually obtained by weighting and summing factors such as information gain, data scarcity attribute, and model performance improvement attribute. The data weight of the same category refers to the overall influence coefficient of that category of data in training sampling, loss calculation, or resource allocation. Data weight is used to control the degree to which the target model values that category of data.
[0051] Specifically, the data summary information can be grouped by category first. For each data summary information within the same category, its information gain, data rarity attribute, and model performance improvement attribute are extracted. These attributes are then merged into a single data contribution attribute using a preset method. The contribution attributes of all data summary information within the same category are summed or aggregated to obtain the cumulative contribution attribute for that category. After obtaining the cumulative contribution attribute for the same category, it can be used as the basis for adjustment. If the cumulative contribution attribute of a category is high, the data weight corresponding to the data summary information of that category can be increased, giving it a higher proportion in subsequent training or evaluation; if the cumulative contribution attribute of a category is low, the data weight corresponding to the data summary information of that category can be decreased, reducing its impact in subsequent training or evaluation, and allowing more training resources to be allocated to data that is more valuable to the target model.
[0052] Optionally, for the same data provider, if the similarity of multiple first material data provided by the data provider is higher than a preset similarity threshold, then the data provider is marked as a first label; for the first material data provided by the first labeled data provider, a preset weight is assigned to the first material data, so as to adjust the contribution attribute of the first material data based on the preset weight; The contribution attribute is used to characterize the resource allocation information corresponding to the first material data.
[0053] It should be noted that similarity is a numerical indicator used to measure the degree of similarity between two or more primary source data. The higher the similarity value, the more similar the primary source data. The preset similarity threshold is a pre-defined similarity judgment standard. The first label refers to the identifier set by the system for a data provider when the similarity of multiple primary source data submitted by that data provider exceeds the preset similarity threshold. For data providers with the first label, a pre-defined differentiated processing strategy is applied to their submitted data in subsequent processing. The preset weight is a pre-defined weighting coefficient used to amplify or reduce the influence of primary source data when calculating contribution attributes. Contribution attributes are indicator values used to quantify the value or contribution of primary source data, serving as the basis for resource allocation. Resource allocation information refers to the determined resource allocation results or rule parameters, such as the number of rewards, resource quotas, or priorities.
[0054] Specifically, for multiple pieces of primary source material data submitted by the same data provider, the similarity between these pieces of primary source material data is calculated. If the similarity exceeds a preset similarity threshold, the primary source material data submitted by that data provider is considered to have high repetition. Based on this, the data provider is given a first label for subsequent differentiation processing. The primary source material data submitted by the first-labeled data provider is processed, and preset weights are assigned to these pieces of primary source material data as weighting parameters in subsequent calculations. When calculating the contribution attribute of each piece of primary source material data, the preset weights are introduced to adjust the contribution attribute. The adjusted contribution attribute is used to characterize the resource allocation information ultimately corresponding to that primary source material data. Based on the contribution attribute, the resource allocation result that the primary source material data can obtain can be determined.
[0055] In this embodiment, in response to the received first material data satisfying the preset condition of a normal distribution, the contribution attribute corresponding to the data provider of the first identifier is updated.
[0056] It should be noted that the data volume of the first source data refers to the quantity or scale of the first source data received within the statistical period. For example, the data volume of the first source data can be the number of data entries or the number of samples. The preset condition of satisfying the normal distribution means that the data volume meets the system's pre-set judgment criteria, so that the distribution of the statistics calculated based on these samples can be approximately regarded as a normal distribution.
[0057] Specifically, the system continuously receives initial data and calculates its volume. When the cumulative sample size meets a preset condition that the relevant statistical results are approximately normally distributed, an update action is triggered. For data providers with the first identifier, their corresponding contribution attributes are updated. This update typically involves correcting the data provider's contribution attributes with more stable and reliable statistical estimates for subsequent resource allocation or evaluation.
[0058] Optionally, based on the cumulative contribution attribute corresponding to the data summary information, the first resource of the data provider to which the data summary information belongs is determined; and the first resource is distributed to the target account corresponding to the data provider.
[0059] It's important to note that the first step is to obtain the cumulative contribution attribute for each data summary, thus determining the overall contribution of this batch of data to the target model. Based on this cumulative contribution attribute, the primary resource that the data provider should receive for that data summary is calculated and determined. This could be a reward amount, computing power quota, points, or traffic support. Finally, the determined primary resource is distributed to the target account linked to the data provider.
[0060] The technical solution of this disclosure first dynamically adjusts the weights of multiple evaluation indicators using the model performance improvement attribute in the indicator data. Then, for data summary information of the same category, it calculates the cumulative contribution attribute of the category by comprehensively considering information gain, data rarity, and model performance improvement attribute, and adjusts the data weight of the category accordingly. Simultaneously, it performs similarity detection on multiple first material data submitted by the same data provider. If the similarity exceeds a preset similarity threshold, the data provider is first-marked, and a preset weight is assigned to the first material data it provides to adjust its contribution attribute. When the amount of received first material data meets the preset condition of a normal distribution, the contribution attribute of the first-marked data provider is updated. This can improve the accuracy of data value assessment and resource allocation efficiency, suppress inflated contributions and resource abuse caused by duplicate data, and update the contribution attribute only when the sample size meets the condition of an approximate normal distribution, thus improving the statistical stability and fairness of contribution assessment.
[0061] Example 3 As an optional embodiment of the present invention, an example is provided to further illustrate the invention.
[0062] It should be noted that high-quality and traceable training materials are extracted from multi-source heterogeneous data, and standardized to ensure their suitability for large-scale model training. The data covers multiple domains and languages, and its sources include, but are not limited to, user-uploaded text, audio, images, and videos. To ensure data legitimacy and traceability, upon receiving the raw data, a hash fingerprint is first generated. The SHA-256 algorithm is used to generate a unique hash value for each data entry, and this hash value is bound to the data provider's identity information and stored in an off-chain database, forming a data traceability record.
[0063] The next stage is data cleaning, which combines rule-based regular expression matching with natural language processing techniques to remove noisy data. This includes removing duplicate content, illegal characters, and advertising information. For text data, TF-IDF and word vectors such as Word2Vec or BERT are used for semantic deduplication. For image or video data, CNN feature extraction and similarity comparison algorithms such as Cosine Similarity are used for deduplication.
[0064] A semi-automatic annotation mechanism is introduced in the data labeling stage, combining pre-trained models such as BERT or RoBERTa for initial classification and label generation, followed by manual review to ensure an accuracy rate of no less than 95%. Formatting is then performed according to model input requirements, converting the data to JSON, CSV, or HDF5 formats and dividing it into training, validation, and test sets in batches. This process ensures data integrity, consistency, and usability, providing a high-quality input foundation for the subsequent construction of the value assessment model.
[0065] The core architecture of this invention is divided into five layers: a data access layer, a value assessment layer, a blockchain interaction layer, an application interface layer, and a governance and arbitration layer. These layers communicate and exchange data through standardized interfaces. The data access layer is responsible for receiving and parsing multimodal data, supporting input in various formats such as text, audio, and images. The data access layer has a built-in multimodal parser that can automatically identify data types and call the corresponding preprocessing modules. Simultaneously, the data access layer integrates a zero-knowledge data capsule generator to generate data summaries that can be used for evaluation without exposing the original plaintext data.
[0066] The value model calculation scheme is as follows: Increment the performance of the validation set... View it as a local manifold drift vector in a high-dimensional representation space ,if Compared with the historical average drift direction The included angle If the size decreases, the weight of the data block in this round is reduced. If If the value model suddenly shifts to a new direction, the reward for parallelism in this round of weighting will be increased. This is because the value model only cares about the performance increment on the validation set. The contribution, not how much the training loss decreased. Therefore, first... Backpropagation is performed to the sample level to obtain the loss score for each sample. The influence function is then applied to the samples... An approximate leave-one-out calculation is performed: the final value model first treats the entire batch of data as a single direction to determine the total weight, and then uses the influence / credit method to reverse the total weight back to the samples. The best samples receive the vast majority of the scores, while the worst samples receive close to zero.
[0067] The implementation steps are as follows: After each validation sample is processed by the current encoder, the [CLS] vector is extracted to form a set. Use local linear embedding or diffusion map to... Projected onto a 3-d latent space, denoted as ;calculate Compared with the previous manifold Procrustes alignment error Let the weighting factor .in It is Sigmoid. and The meta-parameters are learnable but not updated via gradients; instead, a Bayesian optimization search is performed every k epochs. Final integral = base contribution score × .
[0068] When the same data provider appears N times consecutively When the value is > 0.95, meaning the directions are almost identical, the first flag is triggered: subsequent data from that address has its weight multiplied by 0.1 and does not recover with any performance improvement; it must wait until the data it provides expands the manifold to a new region, such as... The effect will be automatically lifted only if the increase exceeds 2.
[0069] Regarding resource consumption, an asynchronous bypass evaluation scheme is proposed, where all recalculations are moved to the CPU pool or even an offline queue, resulting in almost zero waiting time on the training end. The implementation plan is as follows: The system roles are divided into Trainer access and Side-CarEvaluator. The Trainer is responsible for normal forward-backward parameter updates; every K steps, the latest checkpoint and the list of batch-ids used for training are put into the message queue in a non-blocking manner.
[0070] The Side-Car Evaluator is responsible for subscribing to the queue and completing the following asynchronously: loading the latest checkpoint → inference verification set → obtaining... ([CLS] vector); perform LLE / diffusion map → obtain 3-D manifold ;and Procrustes alignment → ;calculate According to the formula, Write back batch-id → Key-value pairs.
[0071] Before the next round of actual sampling data, the Trainer only needs time to pull the latest data from the KV cache. This is used for weighted Dataloader or gradient scaling. Synchronous build memory update rules: only... Only truly fresh samples are allowed to be used. Write to the memory; if the memory size exceeds the budget, use Reservoir Sampling to randomly evict old samples, or prioritize evict the least used ones. Similarity search uses FAISS-IVF or ScaNN to perform a full similarity search. The calculations are placed in the Side-Car, and as well as Package and output them together. The Side-Car Evaluator calculates... First, we look up a memory, which stores the [CLS] vectors corresponding to all historically high-weighted samples. Then, for each sample in the current batch... Calculate its minimum cosine distance to the memory bank: Then give an exponential decay coefficient. , It is a temperature hyperparameter.
[0072] Multiply the original score obtained from influence / credit directly by Then, normalization ensures that closely related data with slight perturbations in the feature space are significantly downweighted, guaranteeing that the model is always pushed towards unexplored regions. Based on repeated sample data, a repetition penalty mechanism is constructed and incorporated into the weight formula. Among these, The minimum cosine distance to the memory bank is multiplied together with the preceding manifold drift signal to complete the implementation of the penalty mechanism.
[0073] Through the above, all recomputation is bypassed by the asynchronous CPU, and the training end only needs to fetch data, train, and drop it into the queue; the queue + cache + rollback mechanism ensures that the GPU is not blocked, and manifold drift monitoring can be implemented in real time.
[0074] Once the value reward model is complete, a smart contract is deployed in the blockchain interaction layer to automatically execute the distribution and recording of points. This layer supports multiple mainstream blockchain platforms, such as Ethereum, Hyperledger Fabric, and Polygon. The smart contract is written in Solidity and features event logging, point minting, transfer, and destruction functions. Furthermore, this layer encapsulates a cross-chain gateway, supporting point mirroring and batch writing between heterogeneous chains to reduce gas costs.
[0075] It should also be noted that, see Figure 3Data Provider Uploads Materials: The data provider submits the first piece of material data. SHA-256 Fingerprint Generation and Identity Binding: A hash fingerprint, part of the data summary information, is generated for each piece of first-piece material data and bound to the data provider's identity for ownership verification / tracing. Cleaning and Static Evaluation: The individual material is cleaned, deduplicated / detected for similarity, and a basic quality assessment is performed. Base Score Generation: A base contribution score is generated for each piece of material. Assembly Waiting: The individual result enters the waiting queue for assembling data blocks. Data Block Assembly: Several processed materials are grouped into data blocks in batches and input into the large model training process. Performance Monitoring: Performance changes are monitored during training / validation, and performance increment signals related to the current batch are calculated. Data Block Weight Generation: Batch weight factors are generated based on performance increments and directional changes, representing the weight of the current batch of data blocks in effectively improving the model. Final Reward Score Calculation: The individual base score Base is combined with the batch weights to obtain the final reward score for the individual material. Smart Contract: Tokens are minted / distributed based on the final reward score and transferred to the data provider's target account, i.e., token deposit / wallet. The on-chain log records the issuance event and uses it for auditing and tracing. The settlement layer takes over individual parameters / batch parameters and maps off-chain evaluation results into on-chain executable issuance actions.
[0076] The technical solution of this disclosure quantifies the value of data in model training and maps contributions to traceable and tradable on-chain tokens. Rewards are distributed based on contribution to incentivize data providers. Multi-dimensional indicators such as information gain, scarcity assessment, and performance are employed to ensure objective and scientific evaluation. Blockchain smart contracts are used to achieve full transparency and auditability of the reward process.
[0077] Example 4 Figure 4 This is a schematic diagram of the structure of the blockchain-based large-scale resource allocation device provided in the embodiments of this disclosure, such as... Figure 4 As shown, the device includes: a data summary information determination module 310, an indicator data determination module 320, a first resource allocation module 330, and an on-chain storage module 340.
[0078] A data summary information determination module is used to receive first material data sent by at least one data provider, clean and process the first material data to obtain second material data, and determine the data summary information of the second material data. The first material data includes at least text data, audio data, and / or video data, and the second material data includes at least a portion of the data in the first material data. An indicator data determination module is used to batch input the data summary information into a target model and determine the indicator data of each data summary information under multiple evaluation indicators. The evaluation indicators include gain evaluation indicators, rarity evaluation indicators, and performance indicators. The indicator data includes information gain corresponding to the gain evaluation indicator, data rarity attributes corresponding to the rarity evaluation indicator, and model performance improvement attributes corresponding to the performance indicators. A first resource allocation module is used to send the indicator data of the data summary information to a smart contract for multiple data summary information, so that when the model performance improvement attribute reaches a preset attribute threshold, a first resource is allocated to the data provider corresponding to the data summary information. An on-chain storage module is used to store the data summary information, the data provider, the first resource, and the sending address and time information of the data summary information on the blockchain.
[0079] The technical solution of this disclosure firstly receives first source data sent by at least one data provider, cleans and processes the first source data to obtain second source data, and determines the data summary information of the second source data. Then, the data summary information is batch-input into the target model, and the indicator data for each data summary information under multiple evaluation metrics are determined. Further, for multiple data summary information, the indicator data of the data summary information is sent to a smart contract so that when the model performance improvement attribute reaches a preset attribute threshold, a first resource is allocated to the data provider corresponding to the data summary information. Finally, the data summary information, data provider, first resource, and the sending address and time information of the data summary information are stored on the blockchain. This solves the problems in the prior art where, during the training process of large language models, training data sources are wide-ranging but of varying quality, with a large amount of inefficient, invalid, or even biased data being introduced into the training process, resulting in poor training effects, potential illusions or biases, and wasted computing resources. This disclosure embodiment achieves accurate and fair allocation of large model resources through multi-metric evaluation and smart contracts, thereby improving resource utilization and incentive effects.
[0080] Based on the above technical solutions, the data summary information determination module 310 includes a cleaning processing submodule, used to perform data deduplication based on the word vectors of the first material data to obtain the second material data of the text data when the first material data is text data and / or audio data; and / or, when the first material data is video data, to determine the first feature information of the first material data based on a feature extraction model, and to determine whether to use the first material data as the second material data based on the first feature information and the historical feature information of historical material data.
[0081] Based on the above technical solutions, the indicator data determination module 320 includes a data rarity attribute determination submodule and a model performance improvement magnitude attribute determination submodule.
[0082] The data rarity attribute determination submodule is used to input the data summary information into the target model, determine the probability information of the data summary information based on a first function in the target model, determine the information gain based on the probability information, and determine the data rarity attribute based on the frequency of occurrence of the data summary information in the training set. The model performance improvement attribute determination submodule is used to obtain the previously determined historical performance attribute, wherein the historical performance attribute is determined based on the target model obtained in the previous training round after processing the test set; input the data summary information as training data into the target model for training, process the test set based on the target model to obtain the current performance attribute; and determine the model performance improvement attribute based on the current performance attribute and the historical performance attribute.
[0083] Based on the above technical solutions, the device further includes: a data weight adjustment module, used to adjust the weights of multiple evaluation indicators according to the model performance improvement attribute in the indicator data; for data summary information of the same category, to determine the cumulative contribution attribute of the same category according to the information gain of the data summary information, the data rarity attribute and the model performance improvement attribute, so as to adjust the data weights of the same category based on the cumulative contribution attribute.
[0084] Based on the above technical solutions, the device further includes: a contribution attribute adjustment module, used to, for the same data provider, if the similarity of multiple first material data provided by the data provider is higher than a preset similarity threshold, mark the data provider as a first; for the first material data provided by the first marked data provider, assign a preset weight to the first material data, so as to adjust the contribution attribute of the first material data based on the preset weight; wherein, the contribution attribute is used to characterize the resource allocation information corresponding to the first material data.
[0085] Based on the above technical solutions, the device further includes: a contribution attribute update module, used to update the contribution attribute corresponding to the first identified data provider in response to the preset condition that the amount of received first material data satisfies a normal distribution.
[0086] Based on the above technical solutions, the first resource allocation module 330 is further configured to determine the first resource of the data provider to which the data digest information belongs based on the cumulative contribution attribute corresponding to the data digest information; and distribute the first resource to the target account corresponding to the data provider.
[0087] The blockchain-based large model resource allocation device provided in this disclosure can execute the blockchain-based large model resource allocation method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of the method execution.
[0088] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.
[0089] Example 5 Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Refer to the following... Figure 5 It illustrates an electronic device suitable for implementing embodiments of the present disclosure (e.g., Figure 5 The diagram below shows the structure of the terminal device or server 500. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals). Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0090] like Figure 5 As shown, electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 502 or a program loaded from storage device 508 into random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. An edit / output (I / O) interface 505 is also connected to bus 504.
[0091] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0092] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.
[0093] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0094] The electronic device provided in this embodiment and the blockchain-based large-scale resource allocation method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0095] Example 6 This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the blockchain-based large-scale resource allocation method provided in the above embodiments.
[0096] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0097] In some implementations, the server may communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and may interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0098] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0099] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: Receive first material data sent by at least one data provider, clean the first material data to obtain second material data, and determine the data summary information of the second material data, wherein the first material data includes at least text data, audio data and / or video data, and the second material data includes at least a portion of the data in the first material data; The data summary information is input into the target model in batches, and the indicator data of each data summary information under multiple evaluation indicators are determined. The evaluation indicators include gain evaluation indicators, rarity evaluation indicators and performance indicators. The indicator data includes information gain corresponding to the gain evaluation indicator, data rarity attribute corresponding to the rarity evaluation indicator and model performance improvement attribute corresponding to the performance indicator. For multiple data summary information, the indicator data of the data summary information is sent to the smart contract so that when the performance improvement attribute of the model reaches a preset attribute threshold, the data provider corresponding to the data summary information is allocated a first resource. The data digest information, the data provider, the first resource, and the sending address and time information of the data digest information are stored on the blockchain.
[0100] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0101] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0102] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.
[0103] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0104] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0105] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0106] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0107] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A large-scale resource allocation method based on blockchain, characterized in that, include: Receive first material data sent by at least one data provider, clean the first material data to obtain second material data, and determine the data summary information of the second material data, wherein the first material data includes at least text data, audio data and / or video data, and the second material data includes at least a portion of the data in the first material data; The data summary information is input into the target model in batches, and the indicator data of each data summary information under multiple evaluation indicators are determined. The evaluation indicators include gain evaluation indicators, rarity evaluation indicators and performance indicators. The indicator data includes information gain corresponding to the gain evaluation indicator, data rarity attribute corresponding to the rarity evaluation indicator and model performance improvement attribute corresponding to the performance indicator. For multiple data summary information, the indicator data of the data summary information is sent to the smart contract so that when the performance improvement attribute of the model reaches a preset attribute threshold, the data provider corresponding to the data summary information is allocated a first resource. The data digest information, the data provider, the first resource, and the sending address and time information of the data digest information are stored on the blockchain.
2. The method according to claim 1, characterized in that, The step of cleaning the first source data to obtain the second source data includes: When the first source data is text data and / or audio data, deduplication is performed based on the word vectors of the first source data to obtain the second source data of the text data; and / or, When the first material data is video data, the first feature information of the first material data is determined based on the feature extraction model, and based on the first feature information and the historical feature information of the historical material data, it is determined whether to use the first material data as the second material data.
3. The method according to claim 1, characterized in that, The step of batch inputting the data summary information into the target model and determining the indicator data of each data summary information under multiple evaluation indicators includes: The data summary information is input into the target model, and the probability information of the data summary information is determined based on the first function in the target model; Based on the probability information, determine the information gain; The data rarity attribute is determined based on the frequency of occurrence of the data summary information in the training set.
4. The method according to claim 1, characterized in that, The performance improvement attribute of the model is determined based on the following method: Obtain the historical performance attributes determined in the previous training, wherein the historical performance attributes are determined based on the target model obtained in the previous training round after processing the test set; After the data summary information is used as training data and input into the target model for training, the test set is processed based on the target model to obtain the current performance attributes; The performance improvement attribute of the model is determined based on the current performance attribute and the historical performance attribute.
5. The method according to claim 1, characterized in that, The method further includes: Based on the model performance improvement attribute in the aforementioned indicator data, adjust the indicator weights of multiple evaluation indicators. For data summary information of the same category, the cumulative contribution attribute of the same category is determined based on the information gain of the data summary information, the data rarity attribute, and the model performance improvement attribute, so as to adjust the data weight of the same category based on the cumulative contribution attribute.
6. The method according to claim 5, characterized in that, The method further includes: For the same data provider, if the similarity of multiple first material data provided by the data provider is higher than a preset similarity threshold, then the data provider is marked as a first tag. For the first material data provided by the data provider of the first tag, a preset weight is assigned to the first material data, so as to adjust the contribution attribute of the first material data based on the preset weight; The contribution attribute is used to characterize the resource allocation information corresponding to the first material data.
7. The method according to claim 6, characterized in that, The method further includes: In response to the fact that the amount of the first source data received satisfies the preset condition of a normal distribution, the contribution attribute corresponding to the data provider of the first identifier is updated.
8. The method according to claim 1, characterized in that, The allocation of first resources to the data provider corresponding to the data digest information includes: Based on the cumulative contribution attribute corresponding to the data digest information, the first resource of the data provider to which the data digest information belongs is determined; The first resource is distributed to the target account corresponding to the data provider.
9. A large-scale resource allocation device based on blockchain, characterized in that, include: A data summary information determination module is used to receive first material data sent by at least one data provider, clean and process the first material data to obtain second material data, and determine the data summary information of the second material data, wherein the first material data includes at least text data, audio data and / or video data, and the second material data includes at least a portion of the data in the first material data; The indicator data determination module is used to input the data summary information into the target model in batches and determine the indicator data of each data summary information under multiple evaluation indicators. The evaluation indicators include gain evaluation indicators, rarity evaluation indicators and performance indicators. The indicator data includes information gain corresponding to the gain evaluation indicator, data rarity attribute corresponding to the rarity evaluation indicator and model performance improvement attribute corresponding to the performance indicator. The first resource allocation module is used to send the indicator data of the data summary information to the smart contract for multiple data summary information, so as to allocate the first resource to the data provider corresponding to the data summary information when it is determined that the performance improvement attribute of the model reaches a preset attribute threshold. The on-chain storage module is used to store the data digest information, the data provider, the first resource, and the sending address and time information of the data digest information on the blockchain.
10. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When one or more programs are executed by one or more processors, the one or more processors implement the blockchain-based large-scale resource allocation method as described in any one of claims 1-8.