Visual Crowdsensing Data Contribution Evaluation Method and System for Machine Learning
By integrating the evaluation method of image semantic similarity, clarity and repetition in visual group intelligence perceptual data, the problem of insufficient data quality and quantity in machine learning scenarios is solved, and reasonable evaluation and incentives to user contribution are achieved, meeting the diversity and quantity requirements of deep learning.
Patent Information
- Application Number
- CN202110175365.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-09
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2041-02-09
AI Technical Summary
In machine learning scenarios, the visual group intelligence perceptual data submitted by users have quality problems such as image semantic mismatch, blur and replication, resulting in insufficient number of images in the data set and it is difficult to meet the data quality and quantity requirements of deep learning.
A visual group intelligence perceived data contribution evaluation method and system are designed for machine learning. By integrating the data quality coefficient algorithm of multi-label image semantic similarity, image clarity and repetition, comprehensively considering image quality and quantity, calculating data contribution, and users are encouraged to submit high-quality and diverse image data.
This method can reasonably evaluate the visual quality and user contribution of perceived data. On the basis of ensuring data quality, it encourages users to submit more pictures to meet the requirements of machine learning scenarios for image quality and quantity.
Smart Images

Figure CN112990268B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning, and particularly to a method and system for evaluating the contribution degree of visual crowd sensing data for machine learning. Background Art
[0002] With the explosive popularity of wireless communication, sensor technology, and wireless mobile terminal devices, mobile devices such as mobile phones and tablet computers on the market integrate more and more sensors and have more and more powerful computing and sensing capabilities. As a new type of sensing method, crowd sensing is more and more widely applied. A large number of smartphone users obtain localized information (such as location, context, noise, traffic, etc.) through the sensors of mobile devices, and these information can be aggregated in the cloud to provide large-scale sensing and social intelligence mining.
[0003] Mobile devices are equipped with sensors such as cameras, microphones, gyroscopes, and accelerometers, and can complete different types of sensing tasks such as numerical, audio, image, and video. Among these sensing methods, the method of using the built-in camera of mobile devices for sensing has received more and more attention from the academic and industrial circles in recent years. Professor Guo Bin proposed the concept of visual crowd sensing in 2017. Visual crowd sensing is a special form of mobile crowd sensing, which requires users to obtain detailed information of the target of interest in the real world in the form of images or videos. Since images and videos can provide relatively rich information, they have received extensive attention and are widely applied in urban sensing, scene recognition, disaster relief, environmental detection, etc. In recent years, machine learning has been applied to data analysis in various fields, which has also become a major driving force for promoting the application of mobile crowd sensing. Visual crowd sensing has become an important data acquisition method for constructing image datasets.
[0004] At present, visual crowd sensing has become an important way to construct image datasets in machine learning scenarios. However, there are data quality problems such as image semantic mismatch, image blur, and duplicate images in the data submitted by users. In machine learning scenarios, it is also necessary to encourage users to contribute more diverse image data to solve the problem of insufficient number of images in the dataset. Summary of the Invention
[0005] In view of the above problems, the present invention designs a method and system for evaluating the contribution degree of visual crowdsensing data for machine learning application scenarios, proposes a contribution degree evaluation method that simultaneously considers two factors of data quality and quantity, designs a visual crowdsensing data quality coefficient algorithm that fuses multi-label image semantic similarity, image clarity and repetition degree, and on this basis, proposes a data contribution degree calculation method that comprehensively considers image quality and quantity. Experimental results show that this method can reasonably evaluate the visual quality of sensed data and the user contribution degree, and can motivate users to contribute more pictures on the basis of ensuring data quality to meet the quality and quantity requirements of pictures in machine learning scenarios.
[0006] To achieve the above object, the technical solution of the present invention is as follows:
[0007] A method for evaluating the contribution degree of visual crowdsensing data for machine learning includes the following steps:
[0008] Publish a visual crowdsensing task and establish a task model according to the visual crowdsensing task;
[0009] Task participants submit image data and establish an image model according to the image data;
[0010] Based on the task model and the image model, automatically identify image features of the image set data submitted by task participants through an image classification model, and the image features at least include image semantic similarity, clarity, and repetition degree;
[0011] Based on the image features identified by the image classification model, evaluate the image set data and calculate the total score of the image data set quality;
[0012] Conduct a data contribution degree evaluation according to the total score of the image set quality to obtain the contribution degree score of the image set to be detected.
[0013] Preferably, establishing the task model according to the visual crowdsensing task is specifically to define the task model through a seven-tuple, and the seven-tuple is:
[0014] task = <tid, time, site_set, desc, cost, pic_num, vc_constrain_set>, where tid is the task identifier, time is the execution time limit of the task, site_set is the set of location constraint conditions for task execution, desc is the description information of the task, cost is the maximum reward budget provided by the task publisher for this task, pic_num is the minimum number of pictures or videos to be collected, and vc_constrain_set is the constraint attribute of the crowd-sourced sensing task in terms of vision, where vc_semantic is the semantic constraint of the image content, which is a set of tags and weights used to define the image content.
[0015] Preferably, the semantic constraint of the image content can be defined as:
[0016] vc_semantic = {s1, s2, …, s m}
[0017] where m is the number of semantic tags required in the image, and s m uses a binary tuple to define the semantic tag and its weight of the task, s m = <Γ m , w m >, Γ m is the tag sequence of the m-th semantic tag, R is the number of tags, and the tags in a group of tag sequences are in an OR relationship, and w m is the weight of the m-th tag sequence.
[0018] Preferably, an image model is established according to the image data. Specifically, the picture model is defined by a ten-tuple: pic = <pid, tid, wid, img, ts, loc, dir, light, lables, qlty>, where pid is the image identifier; tid is the task identifier; wid is the participant identifier; img is the image content; ts is the timestamp of the image; loc is the location information of the picture, which is the GPS information of the mobile device, including longitude and latitude; dir is the direction information when the picture is taken, which includes the data sensed by the acceleration sensor and gyroscope sensor of the mobile device; light is the data sensed by the light sensor of the mobile device; lables are the semantic tags of the picture content, which reflect the target of the picture content and are generated after automatically identifying the image based on a deep learning classifier; qlty is used to describe the visual quality of the image.
[0019] Automatically identify the image features of the input image to be detected based on the task model and the image model through an image classifier, which specifically includes the following steps:
[0020] Use a convolutional neural network to construct an image semantic recognizer, an image sharpness classifier, and a duplicate image detector respectively;
[0021] Use the image semantic recognizer to perform multi-label semantic feature recognition on the image to be detected;
[0022] Use the image sharpness classifier to perform sharpness feature recognition on the image to be detected;
[0023] Use the duplicate image detector to automatically detect the repeatability features of the image to be detected, and recognize the image features with high similarity generated by copying, rotating, and cropping the original image.
[0024] Preferably, in the automatic detection of the repeatability features of the image to be detected by the duplicate image detector, the ORB algorithm is used to extract low-order features, and at the same time, a convolutional neural network is used to extract high-order features. After fusing the low-order features and high-order features, the repeatability of the two images is calculated.
[0025] Preferably, based on the image features recognized by the image classification model, evaluate the image set data, and calculate the total score of the image data set quality, which specifically includes the following steps:
[0026] a. Calculate the semantic similarity S(I i ) of the i-th image in the image set to be detected:
[0027] Extract the multi-label semantic classification results of the image semantic recognizer, and calculate the semantic distance between the m label sequences of the task semantic constraint
[0028] vc_semantic and the n classification labels of the image. The result is a two-dimensional vector of m*n. Take the maximum value * weight from each label sequence as the semantic similarity of the label. The calculation formula is as follows:
[0029]
[0030]
[0031] where r = 1, 2,...R
[0032] In the formula, is the semantic distance calculation function between label and u j , w i is the semantic label weight defined in the task semantic constraint, c j is the confidence coefficient of automatically classifying the label u j by the image classifier, and c j The calculation formula is as shown in formula (3):
[0033]
[0034] Wherein, q j is the confidence level of the label u automatically classified by the image semantic recognizer; j θ is the semantic distance threshold;
[0035] b. Calculate the clarity score B(I i ) of the i-th image in the image set to be detected:
[0036] After the image I i is automatically recognized by the image clarity classifier, the output result is the clarity category L i of the image I j and its confidence level ε j . Calculate the score of the clarity classification L j according to formula (4). The formula is as follows:
[0037]
[0038] Wherein, H, M, and L are the three categories of high, medium, and low image clarity classified by the image clarity classifier;
[0039] Then, calculate the image clarity score using formula (5). The formula is as follows:
[0040]
[0041] Wherein, g(L j ) is the score corresponding to the classification label, and ε j is the confidence level of the clarity classification L j output by the classifier;
[0042] c. Calculate the duplication score D(I i ) of the i-th image in the image set to be detected:
[0043] Calculate the duplication score of the i-th image in the image set I to be detected. The calculation formula is as follows:
[0044]
[0045] Wherein, N is the number of images in the image set I, and Dup(i, j) is the duplication score of two images. The calculation formula is as follows:
[0046]
[0047] Wherein, sim(I i , I j) is the optimal ratio of feature points output by the repeated image detector. The larger its value, the more similar the two images are, and the lower the corresponding repetition score. E is the image similarity threshold. When sim(I i , I j ) is less than E, it is considered that there is no repetition relationship between I i and I j , and the repetition score is 1;
[0048] d. Calculate the total quality score Q(k) of the image set to be detected. The formula is as follows:
[0049]
[0050] In the formula, ε, δ, and φ are the weights of semantic similarity, clarity, and repetition respectively, and ε + δ + φ = 1, 0 ≤ Q(k) ≤ N.
[0051] Preferably, the calculation formula for the contribution score of the image set to be detected is as follows:
[0052]
[0053] In the formula, g is the gain parameter, and Z is the minimum number constraint of pictures required by the task.
[0054] A visual crowdsensing data contribution evaluation system for machine learning includes a task model, an image model, an image classifier, and a data evaluator. Among them,
[0055] The task model is used to define the basic task information and task constraint conditions in the visual crowdsensing task;
[0056] The image model is used to define the image data submitted by users, and the image data includes image files, image basic information, and image illumination and position context information;
[0057] The image classifier is a machine learning-based classifier used to automatically classify and identify image-related features;
[0058] The data evaluator is used to evaluate the data quality and contribution of the image data set submitted by users. According to the visual constraints required by the task requester, combined with the classification results of the image-related features by the image classifier, calculate the image quality score through three dimensions of image semantic similarity, image clarity, and image repetition, and calculate the final data contribution score using a contribution algorithm that integrates quality and quantity on this basis.
[0059] Preferably, the image classifier includes an image semantic recognizer, an image clarity classifier, and a repeated image detector. Among them,
[0060] The image semantic recognizer is used to recognize the semantics of the scenes and objects in the picture and output semantic labels and their confidence levels;
[0061] The image clarity classifier is used to automatically classify pictures into three types: high, medium, and low;
[0062] The duplicate image detector is used to extract near-duplicate features of the image, extract global high-order features of the image using a convolutional neural network, and fuse them with low-order features extracted using the ORB algorithm to identify image features with high similarity generated by copying, rotating, and cropping the original image.
[0063] Based on the above technical solutions, the beneficial effects of the present invention are as follows: The main problem solved by the present invention is to evaluate the image data submitted by users in the deep learning scenario of visual crowd sensing. The prior art mainly aims to encourage users to submit high-quality photos and mainly evaluates aspects such as the visual quality and image similarity of the images. However, in the deep learning scenario, since the deep learning image dataset requires a sufficient number of data and requires image diversity, that is, it requires users to contribute high-quality images and also allows users to contribute a large number of low-quality pictures. Therefore, the existing methods that use data quality as the evaluation target are not applicable to the application scenarios of constructing deep learning datasets. Different from the existing crowd sensing technologies that use data quality as the evaluation index, the present invention uses user contribution as the evaluation index, fuses three dimensions of multi-label image semantic similarity, image clarity, and repetition degree to calculate the visual crowd sensing data quality, and on this basis, proposes a data contribution calculation method that simultaneously considers image quality and quantity. For images with low quality, the contribution score can be increased by quantity. Experimental results show that this method can reasonably evaluate the visual quality of the sensed data and user contribution, and can encourage users to contribute more pictures on the basis of ensuring data quality to meet the quality and quantity requirements of pictures in the machine learning scenario. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] The following further details the specific implementation manners of the present invention with reference to the accompanying drawings.
[0065] Figure 1 : Flowchart of the method for evaluating the data contribution of visual crowd sensing for machine learning according to the present invention;
[0066] Figure 2 : Examples of images with different categories of clarity in the method for evaluating the data contribution of visual crowd sensing for machine learning according to the present invention;
[0067] Figure 3: In the method for evaluating the contribution degree of visual crowdsensing data for machine learning of the present invention, when the gain parameter g = 6 and the minimum value constraint N of the number of required pictures for the task is 10 and 6 respectively, the data contribution degree score λ k Function curve graph;
[0068] Figure 4 : Calculation result of image semantic similarity in the method for evaluating the contribution degree of visual crowdsensing data for machine learning of the present invention;
[0069] Figure 5 : Comparison result of Pearson correlation coefficients;
[0070] Figure 6 : Comparison result of data quality evaluation coefficients, where a is the comparison result of the data quality evaluation coefficient of User1, b is the comparison result of the data quality evaluation coefficient of User2, c is the comparison result of the data quality evaluation coefficient of User3, and d is the comparison result of the data quality evaluation coefficient of User4;
[0071] Figure 7 : Principle block diagram of the system for evaluating the contribution degree of visual crowdsensing data for machine learning of the present invention. Detailed implementation manners
[0072] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.
[0073] Embodiment 1
[0074] Such as Figures 1 to 6As shown in the figure, the visual crowd-sourcing perception data contribution evaluation method for machine learning of the present invention is a visual crowd-sourcing perception data quality evaluation method including three factors: image semantics, clarity, and redundancy. Different from other image quality evaluation methods, in addition to considering the visual quality of the picture, the evaluation of image semantic similarity is added. The results show that this method can correctly evaluate situations such as content-irrelevant pictures, picture duplication, and blurred images. In view of the characteristics of visual crowd-sourcing in the machine learning application scenario, a data quality evaluation algorithm based on user contribution is designed, considering both the quality and quantity of picture data. This method includes the following steps: obtaining a visual crowd-sourcing perception task and establishing a task model according to the visual crowd-sourcing perception task; obtaining image data and establishing an image model according to the image data; classifying and recognizing the input image set to be detected based on the task model and the image model; evaluating the image set to be detected based on the classification and recognition results, and calculating the total score of the data set quality; performing data contribution evaluation according to the total score of the data set quality to obtain the contribution score of the image set to be detected. The present invention uses three dimensions of image semantic similarity, image clarity, and image redundancy as the main evaluation indicators of data quality to calculate the quality score of each picture. The specific descriptions of image semantic similarity, image clarity, and image redundancy are as follows,
[0075] a. Multi-label image semantic similarity
[0076] When the task requester publishes a visual crowd-sourcing perception task, in addition to the conventional requirements such as the time and location of collecting pictures, requirements for picture content or shooting targets are usually also put forward. However, in order to complete the task, users may submit pictures that are not relevant to the task requirements. To solve this problem, the task requester can set the semantic label of the image content as a constraint condition when publishing the task. When the user submits a picture, the server-side uses a classifier based on a convolutional neural network to identify the image content semantics and calculates the similarity between the image semantics required by the task and the picture semantics submitted by the user.
[0077] For a crowd-sourcing perception task, the semantic label constraint parameters required by the task publisher can be defined as:
[0078] vc_semantic = {s1, s2, …, s m}
[0079] In the formula, m is the number of semantic labels required in the image, s m A binary tuple is used to define the semantic label of the task and its weight, s m = <Γ m , w m >, Γ m is the label name sequence of the mth semantic label, R is the number of label sequences, is an OR relationship, w m is the weight of the m-th semantic label. For example, for the ecological environment visual crowdsensing task, vc_semantic = {<{river, lake, stream, water body waterfall}, 0.6>, {<grassland, prairie, mountain range, canyon}, 0.4>}. After the user submits an image, a multi-label classifier such as an object recognition or scene recognition classifier is used for multi-label classification to classify the top N labels. The label of the k-th image submitted by the user is wp k = {μ1, μ2, …, μ n}, μ n is the classification result output by the classifier, which is a binary tuple μ n = <lable, score>, lable is the image semantic label, and score is the confidence value corresponding to the label. For example, wp k = {<tree, 0.8>, <canyon, 0.7>, <river, 0.4>, <scenery, 0.2>, <sky, 0.6>}.
[0080] Extract the multi-label semantic classification results of the image semantic recognizer and calculate the task semantic constraint
[0081] the semantic distance between the m label sequences of vc_semantic and the n classification labels of the image. The result is a two-dimensional vector of m * n. Take the maximum value * weight from each label sequence as the semantic similarity of the label. The calculation formula is as follows:
[0082]
[0083]
[0084] where r = 1, 2,... R
[0085] In the formula, is the semantic distance calculation function for the label and u j , W i is the semantic label weight defined in the task semantic constraint, c j is the confidence coefficient for automatically classifying the label u j by the image classifier, c j The calculation formula is as shown in formula (3):
[0086]
[0087] In the formula, q j is the confidence for automatically classifying the label u j by the image semantic recognizer, and θ is the semantic distance threshold.
[0088] The semantic similarity of the k-th image is calculated using Formula (2) and Formula (3) as shown in Algorithm 1.
[0089]
[0090]
[0091] b. Image sharpness
[0092] The sharpness of an image is an important indicator to measure the quality of the image. It can correspond well to people's subjective feelings. Low sharpness of the image indicates image blurriness, which affects the data quality of crowdsensing. Therefore, the present invention takes image sharpness as one of the important indicators for evaluating data quality. Convolutional neural network (CNN) has strong advantages in image classification and image feature extraction. The present invention uses CNN to construct an image sharpness classifier. Using publicly available image datasets with high sharpness as the original images, after Gaussian blur processing, they are used as the datasets for deep learning to train the model. The calculation of Gaussian blur is shown in Formula (12):
[0093]
[0094] where σ is the blur radius, and this parameter determines the degree of image blurriness. The larger σ is, the more blurred the image is. In order to better distinguish the degree of image blurriness and improve the classification accuracy, the present invention sets σ of the images in the M (medium) category to 2, and σ of the images in the L (low) category to 5. Examples of the three categories of H (high), M (medium), and L (low) are shown in Figure 2 as shown.
[0095] With the classification labels being H (high), M (medium), and L (low), the corresponding scores are shown in Formula 4:
[0096]
[0097] The k-th image I k The label set L k ={L k1 , L j2 ,..., L kn}, where n is the number of labels, and L kn ∈{H, M, L}, then the image sharpness score is calculated as shown in Formula 5:
[0098]
[0099] where g(L j ) is the score corresponding to the classification label, as shown in Formula (4), and ε jFor the classifier output clarity classification L j Confidence. When constructing the training set, σ is set to 2, 5, and 7 respectively to construct data sets of three categories with high (H), medium (M), and low (L) levels of blurriness. However, due to the diversity of the blurriness of the images uploaded by users, it is difficult to accurately classify them into these three categories. To obtain a reasonable score for the image clarity, the confidence ε output by the classifier is used in the present invention j As the blurriness coefficient.
[0100] c. Image repeatability
[0101] Image repeatability is mainly used to detect whether the user generates pictures with high similarity through operations such as copying, rotating, and cropping the original picture. A convolutional neural network (CNN) is used to extract the global high-order features of the image, and they are fused with the low-order features extracted using ORB. The Euclidean distance is used to complete the similarity calculation. The two image repeatability scores are shown in formula (7):
[0102]
[0103] Where I i , I j ∈I, are two images in the image set I uploaded by the user User K , sim(I i , I j ) is the proportion of the best feature points extracted by CNN+ORB for image I i and I j . The larger its value, the more similar the two images are, and the lower the corresponding repeatability score. E is the image similarity threshold. When sim(I i , I j ) is less than h, it is considered that there is no repeat relationship between I i and I j , and the repeatability score is 1. The image repeatability of the image set I submitted by the user User K is calculated as shown in formula (6):
[0104]
[0105] d. Data contribution evaluation
[0106] Based on the image quality evaluation, aiming at the requirements of the machine learning image data set for the quantity and diversity of pictures, a contribution evaluation method that simultaneously considers the picture quality and quantity is designed
[0107] For the image set I={I1, I2,..., I K uploaded by the user User NEach picture in} is scored using three metrics: image semantic similarity, clarity, and duplication. The weighted average is taken as the data quality score of the $i$-th picture, and the total score $Q(k)$ of the picture set $I$ k is shown in Equation (1).
[0108]
[0109] where $\epsilon$, $\delta$, $\varphi$ are weights, and $\epsilon+\delta+\varphi = 1$, $0\leq Q(k)\leq N$.
[0110] The contribution of a user is related not only to data quality but also to data quantity. Especially for visual crowd sensing in machine learning scenarios, the sensing task requires users to submit a large number of pictures. To consider the influence of both data quality and quantity on user contribution, based on the total quality score $Q(k)$ of the picture set uploaded by user UserK, the sigmoid function is used to normalize the image quality score. Through the sigmoid function, the scores of each user are distributed between [0, 1]. When $x$ is less than the threshold, $y$ is a concave function, and the increasing rate of the $y$ value is relatively high. This enables the quality score to increase rapidly with the increase in the number and quality of pictures when the data quality is low. When $x$ is greater than the threshold, $y$ is a convex function, and the increasing rate of the $y$ value slows down. This makes it more difficult for users to further improve the quality score when the data quality is already high, and they need to make greater efforts to increase the number and quality of pictures. In the application scenario of machine learning, the requirement for quality is that most pictures are of high quality, but pictures of different qualities are also needed. The deep learning dataset requires a large number of pictures. Therefore, on the premise of minimizing the cost, the maximization of the number of pictures should also be achieved. The sigmoid function can better meet the characteristics of picture data required by machine learning. After modifying the sigmoid function, Equation (9) is used to calculate the data contribution score $\lambda$ that considers both quantity and quality k .
[0111]
[0112] where $g$ is the gain parameter, $Z$ is the minimum constraint of the number of pictures required by the task, $Q(k)$ is the total data quality score of user UserK, and is the total score of the image set uploaded by the user calculated using three dimensions: image semantic relevance ($S$), image clarity ($B$), and image duplication ($D$). The calculation formula is as shown in Equation (1). Let $g = 6$. After the user uploads $N$ high-quality pictures, $\lambda$ k = 0.9975, and the full score can be obtained. After uploading $N / 2$ high-quality pictures, $\lambda$ k = 0.5, and half of the full score can be obtained. When uploading 0 pictures, $\lambda$ k = 0.002, approximately equal to 0 points. When the values of $N$ are 10 and 6 respectively, $\lambda$ kThe function curve of Figure 3 is shown as follows. The value of Q(k) is proportional to the number of pictures. However, since the sigmoid function is used for the final quality score, when the value of Q(k) is large, the quality score increases slowly and does not exceed 1. Users can approach the full score by submitting N high-quality pictures, but the score can still be improved by increasing the number of pictures, thus motivating users to submit more pictures.
[0113] Experimental Results and Analysis
[0114] 1.1. Perception Task
[0115] Design an ecological environment visual crowd-sensing task, the task is to collect water quality observation data of lakes, rivers, etc., including GPS of each observation point and water quality observation photos. According to the photos containing water environment (streams, rivers, lakes, etc.) and elements related to the ecological environment such as plants, the image semantic constraints required by the task requester are as follows:
[0116] vc_semantic = [{'labels': ['river', 'lake','stream', 'water', 'waterfall', 'pond', ”], 'weight': 0.8}, {'labels': ['mountain','meadow', 'tree', 'plant', 'valley'],'weight': 0.2}]
[0117] 1.2. Evaluation Metrics
[0118] To correctly evaluate the semantic similarity algorithm of the present invention, the algorithm of the present invention and the following two methods are used for comparative analysis:
[0119] Directly use the word semantic similarity calculation API of OpenHowNet (Algorithm 2). For each of the Top5 labels identified by the classifier for the pictures uploaded by the user and the image labels required by the task, directly use OpenHowNet to calculate the maximum similarity between the two words based on the concept knowledge base defined in HowNet.
[0120] Do not consider the image semantic correlation index (Algorithm 3), that is, the similarity score is close to 1. For the calculation of the Pearson correlation coefficient, a random number between 0.9 and 1 is taken.
[0121] The Pearson correlation coefficient is used as the evaluation criterion for the quality of the correlation calculation algorithm. Five volunteers rated the relevance of the pictures uploaded by five users according to the task requirements. The rating value ranges from 1 to 5 points, with 5 being the best, 4 being better, 3 being good, 2 being average, and 1 being poor. After normalization, the rating is converted into a value between 0 and 1. The Pearson correlation coefficients are calculated respectively for the similarity scores calculated by three methods and the manual rating values. The larger the value, the better the correlation. The calculation of the Pearson correlation coefficient is as shown in formula (10):
[0122]
[0123] where x is the similarity score vector calculated by three methods, and y is the manual rating of similarity.
[0124] 1.3. Experimental Results and Analysis
[0125] Five users each uploaded 10 photos. The semantic similarity scores of each photo and the average score of the photo set uploaded by the users were calculated. The calculation results are shown in Table 1. Figure 4 The pictures uploaded by User1 and their calculated relevance scores are shown. S is the similarity score calculated by the algorithm of the present invention, and M is the manual evaluation score. Among the 10 pictures uploaded by User1, only the last 2 pictures have a relatively low relevance to the image requirements of the perception task, and the remaining 8 pictures meet the requirements well. The similarity calculated by the algorithm of the present invention is basically consistent with the manual rating.
[0126] Table 1 Image Semantic Similarity Scores
[0127]
[0128] The Pearson correlation coefficient values of each algorithm are shown in Table 2. Through Figure 5 It can be seen that among the three algorithms, the algorithm of the present invention has the highest accuracy in relevance calculation. Compared with Algorithm 2, the accuracy is significantly improved after adding weights and image classification confidence calculation. Algorithm 3 has the worst relevance accuracy, indicating that if the relevance calculation is not performed on the images uploaded by users, the determination accuracy of the quality of crowd-sourced perception data will be greatly reduced, and it will also have a greater impact on the accuracy of the user reward calculation.
[0129] Table 2 Pearson Correlation Coefficient between Image Semantic Similarity and Manual Rating
[0130]
[0131] 2. Image Clarity
[0132] 3,750 images were selected from the NUS-wide-128 dataset and processed using the Gaussian blur method of Formula 4, with 1,250 images for each category. Transfer learning was performed using the pre-trained image classification model of EasyDL, and the accuracy and recall rate of the trained image clarity classification model reached 99.7% and 99.6% respectively.
[0133] 3. Image Duplication
[0134] Five groups of images were selected from the NUS-wide-128 dataset to simulate the image sets uploaded by 5 users, with 10 images for each user. The images were rotated, cropped, and color-adjusted to simulate the replication operations of the images. The data and replication score calculations are shown in the table. Normal is the number of normal images, AbNormal are the abnormal images with replication operations such as rotation and cropping, GT is the manual scoring value after normal images are scored 1 point and abnormal images are scored 0 point, D is the image duplication score of the algorithm of the present invention, and ACC is the accuracy. It can be seen from Table 3 that the algorithm of the present invention can correctly identify images with a replication relationship and can obtain a relatively accurate score.
[0135] Table 3 Comparison of Image Duplication Scoring Performance
[0136]
[0137] 4. Data Contribution Score
[0138] 4.1. Experimental Setup
[0139] To evaluate the performance of the data contribution evaluation method based on Formula (10), 4 groups of images were randomly selected from the NUS-wide-128 dataset, and operations such as blurring, copying, cropping, and stretching were performed on the images to simulate the image sets submitted by 4 users. The data quality evaluation coefficient of each image was calculated using Formula (10), and the number of images in the constructed test dataset is shown in Table 4.
[0140] Table 4 Data Quality Evaluation Test Dataset
[0141]
[0142] To evaluate the performance of the data quality evaluation coefficient based on contribution degree proposed by the present invention (Contribution method of the algorithm of the present invention), a data quality evaluation coefficient based on the average value (Mean algorithm) as shown in Formula (11) was designed for comparative analysis.
[0143]
[0144] Among them, Z is the minimum number of image constraints required by the task, m is the number of images submitted by the user, and Q(k) is the total data quality score calculated by user k according to formula (1).
[0145] 4.2. Experimental Results
[0146] For the 4 image sets set in 5.4.1, two methods are used to calculate the data quality evaluation coefficient respectively. The minimum number of image constraints N required by the task is 10, and the slope parameter g of the Contribution method is 6. The calculation results are as Figure 6 shown. User1 submitted a total of 25 images, but the quality of the images is not high. Through Figure 6 a, it can be seen that when using the Mean method, when the number of images reaches 10, due to the submission of low-quality images, the data quality evaluation coefficient of the entire image set will be affected by the average value and decrease, which will affect the enthusiasm of users to submit more images; when using the Contribution method, although the evaluation coefficient of the 10 images submitted by the user is not high, the evaluation coefficient can be improved by submitting more images. When the user submits 25 images, the evaluation system is close to 1.0, thus motivating users to submit more images; through Figure 6 b, it can be seen that when the evaluation coefficient of user2 is between 0.2 and 0.8, the Contribution method can quickly improve the evaluation coefficient by increasing the number of higher-quality images, and the improvement rate is much faster than the Mean method; Figure 6 c shows that the quality of the first 10 images submitted by user3 is very high, and the quality of the 11th and 12th images slightly decreases. If the Mean method is used, the evaluation coefficient will decrease, forcing the user to delete the 11th and 12th images, but the Contribution method makes the evaluation coefficient increase slowly, and the user will submit the 11th and 12th images; Figure 6 d shows that the quality of the 10 images submitted by user4 is relatively low, and the evaluation coefficient of the Contribution method is lower than that of the Mean method, so before the minimum number of image requirements is reached, it plays a better role in restricting low-quality images.
[0147] Through the comparative analysis of the experimental results of the above 4 image sets, it can be seen that the evaluation method based on contribution degree adopted by the present invention pays attention to both the quality and quantity of images at the same time. Users can obtain a higher score through high-quality images, or can improve the score by submitting more images of different qualities.
[0148] Such as Figure 7As shown in the figure, the visual crowd-sourcing perception data contribution evaluation system for machine learning consists of four parts: a task module, an image model, an image classifier, and a data evaluator. Although the visual crowd-sourcing perception task mainly focuses on image data perception, the perception tasks also have diversity. In order to define tasks with different types of requirements and constraints, a flexible multi-task model needs to be established. The visual crowd-sourcing system model of the present invention consists of a task model and an image model, which are specifically described as follows:
[0149] 1. Task Model
[0150] In order to collect highly relevant data, the task publisher needs to define multi-dimensional constraints for the acquisition and quality evaluation of pictures. The definition of the task should use quantifiable parameters to guide picture acquisition. Therefore, the definition of the visual crowd-sourcing perception task consists of multiple elements such as time, location, and the number of pictures. A task can be expressed as a seven-tuple:
[0151] tsk = <tid, time, site_set, desc, cost, pic_num, vc_constrain_set>
[0152] Among them, tid is the task identifier, time is the execution time limit of the task, including the start time and end time of the task. site_set is the set of location constraint conditions for executing the task. A task can have multiple locations, and each location is defined as a three-tuple <longitude, latitude, radius>, which respectively represent the longitude and latitude of the center point and the radius. desc is the description information of the task, cost is the maximum reward budget provided by the task publisher for this task, and pic_num is the minimum number of pictures or videos to be collected. The number of task participants is not defined here. The number of participants is determined by the server-side system when recruiting participants with the goal of maximizing utility based on these two parameters, cost and pic_num, as well as information such as the reputation value and quote of the participants.
[0153] The above parameters are common parameters for mobile crowd sensing tasks. In the visual crowd sensing that this invention focuses on, vc_constrain_set is a parameter that plays a key role in data quality. It represents the set of constraints for the sensing task in terms of vision. To meet the requirements of image diversity in the machine learning scenario, the visual constraints of the task can be flexibly defined. Here are several frequently used constraint parameters: vc_g is the threshold of geographical distance, and the sensed data within the range of vc_g will be regarded as redundant; vc_a is the multi-view shooting angle constraint for the same target; vc_light is the environmental light intensity constraint for the acquired pictures; vc_sim is the image similarity threshold, and those with similarity exceeding this value will be deleted to reduce data redundancy; vc_blur is the image sharpness parameter, used to constrain the sharpness of the image; vc_s is the semantic parameter of the image content, which is a set of labels and weight values used to define the image content. The three parameters vc_sim, vc_blur, and vc_semantic are particularly important for data quality control in the machine learning scenario.
[0154] 2. Image Model
[0155] In the visual crowd sensing task, the sensed data is mainly pictures or videos. When participants take photos, the mobile device will simultaneously record the context information in addition to the image. To represent the image and its context information, this invention defines a picture model through a ten-tuple:
[0156] pic = <pid, tid, wid, img, ts, loc, dir, light, lables, qlty>
[0157] Among them, pid is the picture identifier, tid is the task identifier, wid is the participant identifier, img is the image content, ts is the timestamp of the picture, and these elements describe the basic information of the picture. loc, dir, and light describe the context information of the picture. loc is the location information of the picture, which is the GPS information of the mobile device, including longitude and latitude; dir is the direction information when the picture is taken, which includes the data sensed by the acceleration sensor and gyroscope sensor of the mobile device, and the shooting direction of the picture can be calculated; light is the data sensed by the light sensor of the mobile device, reflecting the environmental light intensity of the picture. lables are the semantic labels of the picture content, reflecting the target of the picture content. It is a special element set by this invention for visual crowd sensing in the machine learning scenario and is generated after automatically identifying the image based on a deep learning classifier. It is a key parameter for calculating the image semantic matching degree in this invention. qlty is used to describe the visual quality of the picture and is also generated after automatically identifying the image quality through a deep learning classifier in this invention.
[0158] 3. Image Classifier
[0159] The image classifier is various machine learning-based classifiers used for automatically classifying and recognizing image-related features. Among them, the image semantic recognizer is used to recognize the semantics of the scenes and objects in the picture, and output semantic labels and their confidence levels. The image clarity classifier is used to automatically classify pictures into three types: high, medium, and low. The complexity detector is used to extract nearly repetitive features of the image.
[0160] 4. Data Evaluator
[0161] The data evaluator is used to evaluate the data quality and contribution degree of the image dataset submitted by the user. According to the visual constraints required by the task requester, combined with the classification results of the image classifier for image-related features, it calculates the image quality score through three dimensions: image semantic similarity, image clarity, and image repeatability. On this basis, it uses a contribution degree algorithm that combines quality and quantity to calculate the final data contribution degree score.
[0162] The above is only the preferred implementation mode of the method and system for evaluating the data contribution degree of vision crowdsourcing perception oriented to machine learning disclosed in the present invention, and is not used to limit the protection scope of the embodiments of this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the embodiments of this specification shall be included within the protection scope of the embodiments of this specification.
[0163] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0164] It should also be noted that the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or elements inherent to such process, method, commodity, or device. Without further limitations, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, commodity, or device including the said element.
[0165] The embodiments in this specification are all described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiments.
Claims
1. A method for evaluating the contribution degree of visual crowd-sensing data for machine learning, characterized in that, It includes the following steps: Release a visual crowd-sensing task and establish a task model according to the visual crowd-sensing task; Task participants submit image data and establish an image model according to the image data; Based on the task model and the image model, automatically identify image features of the image set data submitted by task participants through an image classification model. The image features at least include image semantic similarity, clarity, and repeatability; Based on the image features identified by the image classification model, evaluate the image set data and calculate the total quality score of the image data set, which specifically includes the following steps: a. Calculate the semantic similarity S(I i ) of the i-th image in the image set to be detected: Take out the multi-label semantic classification results of the image semantic recognizer and calculate the task semantic constraint The semantic distance between the m label sequences of vc_semantic and the n classification labels of the image, and the result is a two-dimensional vector of m*n. Take the maximum value * weight from each label sequence as the semantic similarity of the label, which is used to define a set of labels and weight sets of the image content. The calculation formula is as follows: where r = 1, 2,... R In the formula, is the label and the semantic distance calculation function of u j , w i is the semantic label weight defined in the task semantic constraint, c j is the confidence coefficient of automatically classifying the label u j by the image classifier, and the calculation formula of c j is shown in formula (3): where q j is the confidence of the label u j automatically classified by the image semantic recognizer, and θ is the semantic distance threshold; b. Calculate the sharpness score B(I i ) of the i-th image in the image set to be detected: Image I i After being automatically recognized by the image sharpness classifier, the output result is the sharpness category L i of Image I j and its confidence level ε j , calculate the score of the sharpness classification L j according to formula (4), and the formula is as follows: In the formula, H, M, and L are the three categories of high, medium, and low image clarity classified by the image clarity classifier; Then, calculate the image clarity score using formula (5), and the formula is as follows: where g(L j ) is the score corresponding to the classification label, and ε j is the confidence of the classifier output for the clarity classification L j ; c. Calculate the repeatability score D(I i ) of the i-th image in the image set to be detected: Calculate the repeatability score of the i-th image in the image set I to be detected, and the calculation formula is as follows: In the formula, N is the number of images in the image set I, and Dup(i, j) is the repeatability score of two images. The calculation formula is as follows: wherein, sim(I i , I j ) is the optimal feature point ratio output by the duplicate image detector. The larger its value, the more similar the two images are, and the lower the corresponding duplication score. E is the image similarity threshold. When sim(I i , I j ) is less than E, it is considered that there is no duplicate relationship between I i and I j , and the duplication score is 1; d. Calculate the total quality score Q(k) of the image set to be detected. The formula is as follows: where ε, δ, are the weights of semantic similarity, clarity, and repetition respectively, and 0 ≦ Q(k) ≦ N; Conduct a data contribution evaluation based on the total quality score of the image set to obtain the contribution score of the image set I to be detected. The formula is as follows: In the formula, g is the gain parameter, and Z is the minimum number constraint of pictures required by the task.
2. The method for evaluating the contribution degree of visual crowd-sensing data for machine learning according to claim 1, characterized in that, The task model established according to the visual crowd-sensing task is specifically defined by a seven-tuple. The seven-tuple is: task = <tid, time, site_set, desc, cost, pic_num, vc_constrain_set> Among them, tid is the task identifier, time is the execution time limit of the task, site_set is the set of location constraint conditions for executing the task, desc is the description information of the task, cost is the maximum reward budget provided by the task publisher for this task, pic_num is the minimum number of pictures or videos to be collected, and vc_constrain_set is the constraint attribute of the crowd-sensing task in terms of vision.
3. The method for evaluating the contribution degree of visual crowd-sensing data for machine learning according to claim 2, characterized in that, The task semantic constraint can be defined as: vc_semantic = {s1, s2, …, s m} where m is the number of semantic labels required in the image, s m A binary tuple is used to define the semantic label of the task and its weight, s m = <Γ m , w m >, Γ m is the label sequence of the m-th semantic label R is the number of tags. Tags in a group of tag sequences have an OR relationship, w m is the weight of the m-th label sequence.
4. The method for evaluating the contribution degree of visual crowdsensing data for machine learning according to claim 1, wherein, The image model established according to the image data is specifically defined by a ten-tuple. The ten-tuple is: pic = <pid, tid, wid, img, ts, loc, dir, light, lables, qlty> Among them, pid is the image identifier; tid is the task identifier; wid is the participant identifier; img is the image content; ts is the timestamp of the image; loc is the location information of the picture, which is the GPS information of the mobile device and includes longitude and latitude; dir is the direction information when the picture is taken, which includes the data sensed by the acceleration sensor and gyroscope sensor of the mobile device; light is the data sensed by the light sensor of the mobile device; lables are the semantic labels of the picture content, which reflect the target of the picture content and are generated after automatically identifying the image based on a deep learning classifier; qlty is used to describe the visual quality of the image.
5. The method for evaluating the contribution degree of visual crowdsensing data for machine learning according to claim 1, wherein, Based on the task model and the image model, automatically identify the image features of the input image to be detected through an image classifier, specifically including the following steps: Use a convolutional neural network to construct an image semantic recognizer, an image sharpness classifier, and a duplicate image detector respectively; Use the image semantic recognizer to perform multi-label semantic feature recognition on the image to be detected; Use the image sharpness classifier to perform sharpness feature recognition on the image to be detected; Use the duplicate image detector to automatically detect the repeatability features of the image to be detected, and identify the image features with high similarity generated by copying, rotating, and cropping the original image.
6. The method for evaluating the contribution degree of visual crowdsensing data for machine learning according to claim 5, wherein, In the automatic detection of the repeatability features of the image to be detected by the duplicate image detector, the ORB algorithm is used to extract low-order features, and at the same time, a convolutional neural network is used to extract high-order features. After fusing the low-order features and high-order features, the repeatability of the two images is calculated.
7. The system for evaluating the contribution degree of visual crowdsensing data for machine learning, wherein, It includes a task model, an image model, an image classifier, and a data evaluator. Among them, The task model is used to define the basic task information and task constraint conditions in the visual crowd-sourcing perception task; The image model is used to define the image data submitted by the user, and the image data includes image files, basic image information, and the illumination and location context information of the image; The image classifier is a machine learning-based classifier used to automatically classify and identify image-related features; The data evaluator is used to evaluate the data quality and contribution degree of the image data set submitted by the user. According to the visual constraints required by the task requester, combined with the classification results of the image-related features by the image classifier, calculate the image quality score through three dimensions of image semantic similarity, image sharpness, and image repeatability, and on this basis, use a contribution degree algorithm that combines quality and quantity to calculate the final data contribution degree score. Among them, Calculating the quality score of the image data set specifically includes the following steps: a. Calculate the semantic similarity S(I i ) of the i-th image in the image set to be detected: Take out the multi-label semantic classification results of the image semantic recognizer and calculate the task semantic constraint The semantic distance between the m label sequences of vc_semantic and the n classification labels of the image, and the result is a two-dimensional vector of m*n. Take the maximum value * weight from each label sequence as the semantic similarity of the label, which is used to define a set of labels and weight values of the image content. The calculation formula is as follows: wherer = 1,2,...R In the formula, is the label and the semantic distance calculation function of u j , w i is the semantic label weight defined in the task semantic constraint, c j is the confidence coefficient of automatically classifying the label u j by the image classifier, and the calculation formula of c j is shown in formula (3): where q j is the confidence of the label u j automatically classified by the image semantic recognizer, and θ is the semantic distance threshold; b. Calculate the sharpness score B(I i ) of the i-th image in the image set to be detected: Image I i After being automatically recognized by the image sharpness classifier, the output result is Image I i with the sharpness category L j and its confidence level ε j , calculate the score of the sharpness classification L according to formula (4) j The formula is as follows: In the formula, H, M, and L are the three categories of high, medium, and low image sharpness classified by the image sharpness classifier; Then, the image sharpness score is calculated using formula (5) as follows: where g(L j ) is the score corresponding to the classification label, and ε j is the confidence of the classifier output for the clarity classification L j ; c. Calculate the repeatability score D(I i ) of the i-th image in the image set to be detected: Calculate the repeatability score of the i-th image in the image set I to be detected, and the calculation formula is as follows: Where N is the number of images in the image set I, and Dup(i, j) is the repeatability score of two images. The calculation formula is as follows: where, sim(I i , I j ) is the optimal feature point ratio output by the duplicate image detector. The larger its value, the more similar the two images are, and the lower the corresponding duplication score. E is the image similarity threshold. When sim(I i , I j ) is less than E, it is considered that there is no duplicate relationship between I i and I j , and the duplication score is 1; d. Calculate the total quality score Q(k) of the image set to be detected. The formula is as follows: where ε, δ, are the weights of semantic similarity, clarity, and repetition respectively, and 0 ≦ Q(k) ≦ N; Evaluate the data contribution degree according to the total score of the image set quality, and obtain the contribution degree score of the image set to be detected. The formula is as follows: Where g is the gain parameter and Z is the minimum value constraint of the number of pictures required by the task.
8. The visual crowdsourcing perception data contribution degree evaluation system for machine learning according to claim 7, characterized in that, The image classifier includes an image semantic recognizer, an image sharpness classifier, and a duplicate image detector. Among them, The image semantic recognizer is used to recognize the semantics of the scene and objects in the picture and output the semantic label and its confidence; The image sharpness classifier is used to automatically classify pictures into three types: high, medium, and low; The duplicate image detector is used to extract the near-repeatability features of the image, extract the global high-order features of the image using a convolutional neural network, and fuse them with the low-order features extracted using the ORB algorithm to recognize the image features with high similarity generated by copying, rotating, and cropping the original image.
Citation Information
Patent Citations
Keyword based evaluation expert intelligent search and recommendation method
CN103605665A
Task demand-oriented excitation selection method and terminal
CN109960583A