Truth Mining Method for Multi-Source Unstructured Text Services in the Metaverse Crowdsourcing Environment

By adopting phased processing methods and specific algorithms in the metacosmic crowdsourcing environment, the computational complexity and adaptability problems in multi-source unstructured text data processing are solved, and efficient and accurate truth-mining and personalized intelligent services are achieved.

CN119203046BActive Publication Date: 2025-06-27SOUTHEAST UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411712102.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-27
Publication Date
2025-06-27
Estimated Expiration
2044-11-27

AI Technical Summary

Technical Problem

The prior art has high computational complexity, high modeling and poor adaptability when processing multi-source unstructured text data in a metacosmic crowdsourcing environment. It cannot meet the needs of high concurrency and real-time, and it is difficult to provide personalized and high-quality intelligent services.

Method used

Efficient processing methods are adopted in stages, including semantic preprocessing stage, truth value optimization and feature generation stage, and task clustering and cluster mapping stage. Semantic preprocessing is performed using BERT-based context embedding and KANN-DBSCAN algorithm, truth value optimization is performed through iterative optimization mechanism and adaptive feature generation mechanism, and task clustering is performed using the K-means clustering algorithm.

Benefits of technology

It significantly improves the truth-mining efficiency and accuracy of multi-source unstructured text data, can flexibly adapt to complex and changeable data environments, meet the high concurrency and real-time needs in the metacosmic crowdsourcing environment, and provides personalized and high-quality intelligent services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119203046B_ABST
    Figure CN119203046B_ABST
Patent Text Reader

Abstract

The present invention discloses a true value mining method for multi-source unstructured text services in a metaverse crowdsourcing environment. The present invention proposes a solution in three stages, namely, the semantic preprocessing stage, constructing a high-dimensional content vector representation of the crowdsourcing answers and performing clustering. The true value optimization and feature generation stage, using the optimized true value mining model to evaluate the work quality of the crowdsourcing workers and dynamically generate features. The task clustering and cluster mapping stage, clustering the tasks with generated features, accurately estimating the true category of each task, and finally determining the true value. The present invention improves the accuracy and efficiency of true value mining and expands its applicability in complex unstructured data environments. By integrating multi-source perception data in the metaverse crowdsourcing environment, the present invention can accurately push intelligent services that meet user needs, improving the service quality and user experience of the metaverse.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a truth mining method for multi-source unstructured text services in a metaverse crowdsourcing environment, belonging to the technical fields of the metaverse and crowdsourcing. Background Art

[0002] In recent years, the emergence of crowdsourcing platforms has created a distributed way of problem-solving. Requesters post their problems to the crowdsourcing platform in the form of tasks, and the crowdsourcing platform is responsible for distributing these tasks to workers. Since the data provided by different workers may conflict, in order to provide accurate answers to requesters, the crowdsourcing platform often uses truth mining methods to solve data conflicts and mine the truth.

[0003] With the rise of the metaverse concept, crowdsourcing platforms have been further extended and applied in virtual environments. The metaverse contains a large amount of multi-source perceptual data generated by users, which includes both numerical data and categorical data, and increasingly involves unstructured text data. For example, users may submit text information related to virtual goods, virtual identities, or digital content in the metaverse. In this case, multi-source unstructured text data presents complex and diverse characteristics. Specifically, multi-source unstructured text data contains colloquial expressions, emotional descriptions, and non-standard formats. Traditional truth mining methods usually rely on a single data type or structured data and lack effective processing capabilities for unstructured text data, resulting in limited applicability in the metaverse crowdsourcing environment.

[0004] At the same time, most existing truth mining methods are based on probabilistic graph models, which model different scenarios by setting prior parameters to infer probabilities. However, these methods require a large number of prior parameters to be preset, the model training and inference processes consume huge computing resources, and the computing time increases exponentially with the growth of data volume, unable to meet the high concurrency and real-time requirements in the metaverse crowdsourcing environment.

[0005] In addition, users in the metaverse have an increasing expectation for intelligent services and pursue highly personalized and high-quality experiences. This places higher requirements on truth mining methods, which need to be able to accurately mine the truth of multi-source unstructured text data and generate results that meet user needs. However, due to the fixed prior parameters of the model, the methods based on probabilistic graph models are difficult to flexibly adapt to the diversity and dynamics of user-generated content. At the same time, the performance of these methods varies greatly in different scenarios and datasets, further limiting their ability to meet the high standards of user personalized service requirements in the metaverse environment.

[0006] In view of this, the present invention proposes a method for truth mining of multi-source unstructured text services in a metaverse crowdsourcing environment. This method can solve problems such as high computational complexity, large modeling difficulty, and poor adaptability in the prior art when dealing with unstructured text data. In addition, by collecting and integrating the text information of multi-source perception data in the metaverse crowdsourcing environment and pushing intelligent services that meet user needs, the service quality and user experience in the metaverse crowdsourcing environment can be improved. Summary of the Invention

[0007] In view of the technical problems existing in the prior art, the present invention provides a method for truth mining of multi-source unstructured text services in a metaverse crowdsourcing environment. Through efficient processing in stages and a dynamic adaptive feature generation mechanism, the deficiencies of the prior art are overcome, and the accuracy and efficiency of truth discovery are greatly improved. At the same time, this method can flexibly adapt to complex and changeable data environments and provide strong support for intelligent services in the metaverse.

[0008] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0009] A method for truth mining of multi-source unstructured text services in a metaverse crowdsourcing environment, including the following stages:

[0010] A. Semantic preprocessing stage: Construct a worker set, a task set, and a crowdsourcing answer set based on the data obtained from the metaverse crowdsourcing platform. Use context embedding based on BERT (Bidirectional Encoder Representations from Transformers) to construct a high-dimensional content vector representation of the crowdsourcing answers. Use the KANN-DBSCAN (K-Average Nearest Neighbor Density-Based Spatial Clustering of Applications with Noise) algorithm to perform adaptive clustering on the high-dimensional content vector representation of the answers and assign a class label to each crowdsourcing answer;

[0011] B. Truth optimization and feature generation stage: Construct a truth mining model based on the worker set, the task set, and the set of class labels attached to the crowdsourcing answers. The aim is to minimize the difference between the truth and the class labels of the crowdsourcing answers, and then evaluate the work quality of the workers. Based on the evaluation results, generate features from two dimensions. One dimension is the confidence of the class label, and the other dimension is the average difference of the confidence of the class label;

[0012] C. Task Clustering and Cluster Mapping Phase: For tasks with generative features, the K-means clustering algorithm is used for clustering to obtain clusters equal in number to the number of task categories. Construct the set of category confidence levels for each cluster, and select the category with the highest category confidence level as the category for each task in the cluster. Map each task to its corresponding category one by one to determine the true category of each task, and finally obtain the true answers to the tasks on the metaverse crowdsourcing platform.

[0013] In the above solution, the semantic preprocessing phase can process unstructured text data and overcome the deficiencies of existing technologies in text data processing; the truth value optimization and feature generation phase can flexibly adapt to the diversity and dynamics of user-generated content and overcome the deficiencies of existing technologies in meeting the high-standard requirements of user personalized services; the task clustering and cluster mapping phase can lightweight process a large amount of complex data and overcome the deficiencies of existing technologies in the high concurrency and real-time requirements in the metaverse crowdsourcing environment.

[0014] The specific steps of the semantic preprocessing phase are as follows:

[0015] A1. Define and construct a worker set, a task set, and a crowdsourcing answer set according to the data obtained from the metaverse crowdsourcing platform, where the worker set is represented as , represents the total number of workers, the task set is represented as , represents the total number of tasks, and the crowdsourcing answer set for each worker is represented as , represents the crowdsourcing answer of worker to task . On this basis, the crowdsourcing answer submitted by each worker to each task is represented as a triple , and the triple set ;

[0016] A2. Define the set of high-dimensional content vector representations

[0017] , and use BERT-based context embedding to construct the set of high-dimensional content vector representations

[0018] of the crowdsourcing answer set ;

[0019] A3. Use the KANN-DBSCAN algorithm to perform adaptive clustering on the set of high-dimensional content vector representations of the answers, and assign a category label to each crowdsourcing answer 。

[0020] Step A3 specifically includes the following steps:

[0021] A3.1 Define that for any task there are crowdsourcing answers, and the corresponding set of high-dimensional content vector representations is

[0022] ,

[0023] For any two vectors and , calculate the cosine distance between the two , and construct a distance matrix ,

[0024] A3.2 Sort each row of the distance matrix in descending order;

[0025] A3.3 Define the mean of the distance matrix , and calculate the mean of each column in the distance matrix to generate ones ; ;

[0026] A3.4 Define the minimum number , which is expressed as the mean of the number of vectors whose distance between any two vectors in each column of the distance matrix is less than the matrix mean . Calculate , and generate ones , where is expressed as the set of vectors whose distance between any two vectors in each column of the distance matrix is less than the matrix mean ,

[0027] ,

[0028] A3.5 Traverse each pair of parameters , , and use the algorithm to cluster the data set.

[0029] In A3.5, the algorithm specifically includes the following steps:

[0030] A3.5.1 Randomly select an unvisited vector , and calculate ,

[0031] A3.5.2 If , the vectors and together form a new cluster, and all unvisited vectors in the current cluster are recursively processed in the same way to expand the cluster;

[0032] A3.5.3 If , the vector is a noise vector;

[0033] A3.5.4 For other unvisited vectors, repeat steps A3.5.1 to A3.5.3 until all vectors belong to a certain cluster or are noise vectors;

[0034] A3.5.5 All noise vectors also form a cluster;

[0035] A3.5.6 Return all cluster sets.

[0036] In step B, the truth value optimization and feature generation phase includes the following steps:

[0037] B1. After the semantic preprocessing phase, each worker for each task submits the crowdsourcing answers with a category label , represented as a quadruple , the set of quadruples , defines the set of the working quality of workers , defines and constructs the set of category labels of all tasks submitted by the worker , defines the set of truth values of all tasks , , represents the truth value of the task ;

[0038] B2. Construct a truth value mining model and define the objective function , where the constraint condition is , when , the loss function , when , the loss function ,

[0039] B3. Initialize the working quality of any worker , ,

[0040] B4. Minimize the objective function, and iteratively perform steps B5 and B6 until the objective function converges, to obtain the work quality of each worker and the true values of all tasks ;

[0041] B5. Update the true values ,

[0042] B6. Update the worker quality ,

[0043] B7. Define and construct the category set , denoting the number of categories. For each task there is a corresponding category for any category label , that is , and the set of workers for this task with all category labels is denoted as , ,

[0044] B8. Define the category confidence set for each task , , where denotes the sum of the worker quality for which all crowdsourcing answers for task are ,

[0045] B9. Calculate the category confidence and construct the category confidence set for each task , which is also the feature set for each task , where the category confidence satisfies

[0046] ,

[0047] B10. Add an additional feature to the feature set for each task , , and construct the final feature set .

[0048] In step C, the task clustering and cluster mapping phase includes the following steps:

[0049] C1. Use the K-means clustering algorithm to cluster all tasks with generated features, obtaining K clusters equal to the number of categories,

[0050] C2. Define the set of class confidence levels in the clusters , each cluster 's set of class confidence levels in the cluster

[0051] , ,

[0052] represents the task where the class label in the cluster is ;

[0053] C3. Calculate the set of class confidence levels in the clusters ,

[0054] C4. For each cluster select ,

[0055] , map each task in the cluster to the corresponding class to determine the true class of each task and finally obtain the true answers of the tasks on the metaverse crowdsourcing platform.

[0056] C1. Use the K - means clustering algorithm to cluster all tasks with generative features to obtain clusters equal in number to the number of classes K, which specifically includes the following steps:

[0057] C1.1 Define the clusters in the K - means clustering algorithm , is the number of clusters, set the number of clusters to the number of classes in the tasks , that is ,

[0058] C1.2 Define the cluster centers. For any , select the task with the largest as the initial center of each cluster. If the task has already been selected as a cluster center, then select the next task in descending order as the initial center of the cluster, and so on;

[0059] C1.3 Calculate the Euclidean distance between each task and the initial center of each cluster, and assign each task to the cluster closest to it;

[0060] C1.4 For each cluster, recalculate the center of the cluster. The center of the cluster is the mean of all tasks in the cluster;

[0061] C1.5 Repeat steps C1.3 and C1.4 until the cluster centers no longer change;

[0062] C1.6 Return a cluster.

[0063] An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, a truth mining method for multi-source unstructured text services in the metaverse crowdsourcing environment is implemented.

[0064] A computer-readable storage medium, on which computer instructions are stored, and when the computer instructions are executed by a processor, the truth mining method for multi-source unstructured text services in the metaverse crowdsourcing environment is implemented.

[0065] Advantageous effects: A truth mining method for multi-source unstructured text services in the metaverse crowdsourcing environment provided by the present invention has the following advantageous effects compared with the prior art:

[0066] (1) The present invention realizes efficient truth mining of complex unstructured text data through three stages: a semantic preprocessing stage, a truth optimization and feature generation stage, and a task clustering and cluster mapping stage. This staged processing method not only significantly improves the processing efficiency but also ensures the accuracy and reliability of truth mining. In the highly interactive and complex environment of the metaverse, this staged method can flexibly adapt to multi-source perceptual data and improve the overall service quality;

[0067] (2) In the semantic preprocessing stage, the present invention introduces context embedding based on BERT and the KANN-DBSCAN algorithm. Context embedding based on BERT can deeply mine the semantic information of text data and provide support for the construction of high-dimensional content vectors. The KANN-DBSCAN algorithm clusters the high-dimensional content vectors to accurately identify and isolate noise data. This combined strategy not only efficiently processes unstructured text data but also lays a solid foundation for subsequent truth mining;

[0068] (3) In the truth optimization and feature generation stage, the present invention designs an iterative optimization mechanism and an adaptive feature generation mechanism. The iterative optimization mechanism can dynamically update the worker quality and truth, and the adaptive feature generation mechanism can dynamically generate confidence features of category labels according to the worker's work quality, and combine the average difference features of confidence to enhance the tolerance to fluctuations in worker behavior data, thereby improving the robustness of the system;

[0069] (4) In the task clustering and cluster mapping stage, the present invention utilizes the K-means clustering algorithm to generate clusters equal in number to the number of task categories, and based on the highest value of category confidence, accurately completes the mapping of tasks and categories, ultimately determining the true answers to the tasks on the metaverse crowdsourcing platform. Through a lightweight processing method, this stage efficiently handles massive and complex data, effectively reducing the computational and resource burdens, and overcoming the problems of high concurrency and insufficient real-time requirements faced by existing technologies in the metaverse crowdsourcing environment;

[0070] (5) The method of the present invention not only significantly improves the efficiency and accuracy of truth value mining, but also provides strong data support for intelligent services in the metaverse. By integrating multi-source crowdsourcing data and optimizing data quality, the present invention provides more accurate and reliable basic data for various intelligent applications in the metaverse (such as personalized recommendations, virtual assistants, etc.), not only promoting the innovative development of intelligent services, but also significantly enhancing the user experience and meeting users' needs for highly personalized and high-quality intelligent services. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 is a framework schematic diagram of the method scenario and basic principle of the present invention,

[0072] Figure 2 is a flowchart showing the overall implementation of the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0073] The present invention will be further clarified below in conjunction with the accompanying drawings and specific implementation cases. It should be understood that these examples are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, those skilled in the art will make various equivalent modifications that fall within the scope defined by the appended claims of this application.

[0074] of the present invention.

[0075] Example: As Figure 1 shown, the multi-source unstructured text service truth value mining framework established in the metaverse crowdsourcing environment of the present invention, as Figure 2 shown, the overall implementation process of the method established by the present invention, the specific implementation cases are as follows. The truth value mining method for multi-source unstructured text services in the metaverse crowdsourcing environment includes three stages:

[0076] A. Semantic preprocessing stage:

[0077] A.1 According to the data obtained from the metaverse crowdsourcing platform, as shown in Table 1, define and construct a worker set , a task set , a crowdsourcing answer set . On this basis, the crowdsourcing answer submitted by each worker for each task can be represented as a triple, for example: ,

[0078] Table 1:

[0079]

[0080] A2. Define and construct a set of high-dimensional content vector representations of crowdsourced answers using BERT-based context embeddings ,

[0081] A3. Use the KANN-DBSCAN algorithm to adaptively cluster the set of high-dimensional content vector representations of answers, and assign a class label to each crowdsourced answer , as shown in Table 2;

[0082] Table 2:

[0083]

[0084] B. Ground truth optimization and feature generation phase:

[0085] B1. After the semantic preprocessing phase, each crowdsourced answer submitted by each worker for each task is attached with a class label, which can be represented as a quadruple, for example

[0086] , define and construct the set of worker quality , define and construct the set of class labels for all tasks submitted by each worker , for example, the set of class labels for all tasks submitted by worker , define the set of ground truths for all tasks , , denoted as the ground truth of task ;

[0087] B2. Construct a ground truth mining model and define the objective function , where the constraint condition is , when , the loss function , when , the loss function ,

[0088] B3. Initialize the quality of any worker , ,

[0089] B4. Minimize the objective function and iteratively perform steps B5 and B6 until the objective function converges, and finally obtain

[0090] The work quality of workers and the true value ,

[0091] B5. Update the true value ,

[0092] B6. Update the worker quality ,

[0093] B7. Define and construct a set of categories For each task ,

[0094] There is a corresponding category for each category label, for example , in addition It means that for this task, all category labels are the set of workers, for example, for task , ,

[0095] B8. Define for each task the set of category confidences , , It means that for task the sum of the worker qualities for all crowdsourced answers with the category is ,

[0096] B9. Calculate the category confidence, for example , construct the set of category confidences for each task which is also the set of features for each task , satisfying the set of features , satisfying the set of features , , , ,

[0097] ,

[0098] B10. Add an additional feature to the set of features for each task to construct the final set of features to construct the final set of features ,

[0099] the final set of features , ,

[0100] ,

[0101] ,

[0102] C. Task Clustering and Cluster Mapping Phase

[0103] C1. Use the K-means clustering algorithm to cluster all tasks with generation features to obtain 3 clusters , , ,

[0104] ,

[0105] C2. Define the set of class confidence in the cluster , each cluster 's set of class confidence in the cluster , , represents the tasks whose class labels in the cluster are ;

[0106] C3. Calculate the set of class confidence in the cluster , , , ,

[0107] C4. For each cluster select ,

[0108] , for select , for , select , for , select . Map each task in the cluster to the corresponding class to determine the true class of each task, that is, the final true value. , is mapped to "A", is mapped to "B", , is mapped to "A", that is finally obtain the true answers of the tasks on the metaverse crowdsourcing platform, as shown in Table 3;

[0109] Table 3:

[0110]

[0111] It should be noted that the above embodiments are not intended to limit the protection scope of the present invention. Any equivalent transformation or substitution made on the basis of the above technical solutions falls within the protection scope of the claims of the present invention.

Claims

1. A truth value mining method for multi-source unstructured text services in a metaverse crowdsourcing environment, characterized by: The following steps are involved: A. Semantic preprocessing stage: Based on the data obtained from the Metaverse crowdsourcing platform, we build a set of workers, a set of tasks, and a set of crowdsourcing answers. We use BERT-based context embedding to build a high-dimensional content vector representation of the crowdsourcing answers. We use the KANN-DBSCAN algorithm to adaptively cluster the high-dimensional content vector representation of the crowdsourcing answers and assign a category label to each crowdsourcing answer. B. Truth value optimization and feature generation phase: Build a truth value mining model based on the worker set, task set, and category label set attached to the crowdsourcing answers. C. Task clustering and cluster mapping stage: For tasks with generated features, the K-means clustering algorithm is used for clustering to obtain clusters equal to the number of task categories. A set of category confidences in each cluster is constructed, and the category with the largest category confidence is selected as the category of each task in the cluster. Each task is mapped one by one to its corresponding category, thereby determining the true category of each task and finally obtaining the true answer to the task on the Metaverse crowdsourcing platform; In the semantic preprocessing stage, the specific steps are as follows: A1. Define and construct the worker set, task set, and crowdsourcing answer set based on the data obtained from the Metaverse crowdsourcing platform, where the worker set is represented by S = {s1, s2, ..., s n }, n represents the total number of workers, and the task set is represented by T = {t1, t2, ..., t m }, m represents the total number of tasks, and the crowdsourcing answer set of each worker is expressed as Represents workers i For task t j Based on the crowdsourced answer, each worker s i For each task t j Crowdsourced answers submitted Represented as a triple Triple Set P={P 11 ,P 12 ...,P ij ,...,P nm }; A2. Define a high-dimensional content vector representation set, Building a Crowdsourced Answer Collection Using BERT-Based Contextual Embeddings The high-dimensional content vector representation of D A3. Use the KANN-DBSCAN algorithm to represent the high-dimensional content vector set of the answer Perform adaptive clustering to identify each crowdsourced answer Assign a class label 2. The truth value mining method for multi-source unstructured text services in the metaverse crowdsourcing environment according to claim 1 is characterized in that: Step A3 specifically includes the following steps: A3.1 Definition for any task t j ∈T has n j crowdsourced answers, and the corresponding high-dimensional content vector representation set is For any two vectors and Calculate the cosine distance between the two Construct a distance matrix A3.2 Distance Matrix Sort each row in descending order; A3.3 Defining the distance matrix mean Calculate the distance matrix The distance matrix mean of each column in Generate n j indivual A3.4 Define the minimum number Represented as a distance matrix The distance between any two vectors in each column is less than the matrix mean The mean of the number of vectors, calculated Generate n j indivual in Represented as a distance matrix The distance between any two vectors in each column is less than the matrix mean Vector collection of A3.5 For each pair of parameters 1≤k<n j , traverse, use The algorithm clusters the data set.

3. The truth value mining method for multi-source unstructured text services in the metaverse crowdsourcing environment according to claim 2 is characterized in that: In A3.5, The algorithm specifically includes the following steps: A3.5.1 Randomly select an unvisited vector Calculate A3.5.2 If vector and Together they form a new cluster, and recursively process all unvisited vectors in the current cluster in the same way to expand the cluster; A3.5.3 If vector is the noise vector; A3.5.4 Repeat steps A3.5.1 to A3.5.3 for other vectors that have not been visited until all vectors belong to a cluster or are noise vectors; A3.5.5 All noise vectors also form a cluster; A3.5.6 Return all cluster sets.

4. The truth value mining method for multi-source unstructured text services in the metaverse crowdsourcing environment according to claim 1 is characterized in that: In step B, the truth value optimization and feature generation phase includes the following steps: B1. After the semantic preprocessing stage, each worker s i For each task t j Crowdsourced answers submitted Each comes with a category label Represented as a four-tuple The set of four-tuples P′={P′ 11 , P′ 12 , ..., P′ ij , ..., P′ nm }, define the worker's work quality set W = {w1,w2,...,w i ,...,w n }, define and build workers i The set of category labels for all submitted tasks Define the set of truth values ​​for all tasks Represented as task t j The truth value of B2. Build a truth mining model and define the objective function The constraints are when When when When B3. Initialize the work quality w of any worker i ∈W, B4. Minimize the objective function, iterate steps B5 and B6 until the objective function converges, and obtain the work quality w of each worker. i and the truth value of all tasks B5. Update the true value B6. Update worker quality B7. Define and construct a category set g = {g1, g2, ..., g x ,...,g X }, X represents the number of categories, For each task t j For example, any category label There are corresponding categories X Correspondingly, All category labels for this task are g X The set of workers is denoted as S x , B8. Define each task t j The category confidence set Represents task t j The category of all crowdsourced answers is g X The sum of the quality of workers B9. Calculate category confidence and construct each task t j The category confidence set Each task t j The feature set The category confidence satisfies B10. Give each task t j The feature set Add an additional feature Constructing the final feature set 5. The truth value mining method for multi-source unstructured text services in the metaverse crowdsourcing environment according to claim 1 is characterized in that: In step C, the task clustering and cluster mapping phase includes the following steps: C1. Use the K-means clustering algorithm to cluster all tasks with generated features to obtain clusters equal to the number of categories K. C2. Define the class confidence set θ in the cluster y , each cluster c y The cluster class confidence set Indicates that the class label in the cluster is X; C3. Calculate the set of class confidences in the cluster C4. For each cluster c y Select g x , Each task t in the cluster j Mapped to the corresponding category g x In this way, the true category of each task is determined, and finally the true answer to the task on the Metaverse crowdsourcing platform is obtained.

6. The truth value mining method for multi-source unstructured text services in the metaverse crowdsourcing environment according to claim 5 is characterized in that: C1. Use the K-means clustering algorithm to cluster all tasks with generated features to obtain clusters equal to the number of categories K. Specifically, the following steps are included: C1.1 defines the cluster c in the K-means clustering algorithm as follows: y ,...,c Y }, Y is the number of clusters, set the number of clusters to the number of classes in the task X, that is, Y = X, C1.2 defines cluster centers. For any g x ∈g, choose The biggest task j As the initial center of each cluster, if task t j has been selected as the cluster center, select the next one in descending order The task j as the initial center of the cluster, and so on; C1.3 calculates the Euclidean distance between each task and the initial center of each cluster, and assigns each task to the cluster closest to it; C1.4 For each cluster, recalculate the center of the cluster. The center of the cluster is the mean of all tasks in the cluster. C1.5 Repeat steps C1.3 and C1.4 until the cluster center no longer changes; C1.6 returns K clusters.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, it implements the truth value mining method for multi-source unstructured text services in the metaverse crowdsourcing environment as described in any one of claims 1 to 6 above.

8. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the computer instructions are executed by the processor, the truth value mining method for multi-source unstructured text services in the metaverse crowdsourcing environment as described in any one of claims 1-6 is implemented.

Citation Information

Patent Citations

  • Crowdsourcing truth value inference method based on multiple influence factors

    CN116861235A