A method for ranking mobile application crowdsourcing test reports based on image-text fusion analysis

Through the automated analysis and sorting of crowdsourcing test reports on mobile applications, the problem of manual burden on developers when reviewing reports is solved, and the testing efficiency and accuracy of obtaining defect information is improved.

CN114780373BActive Publication Date: 2025-05-06NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111471921.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-30
Publication Date
2025-05-06
Estimated Expiration
2041-11-30

AI Technical Summary

Technical Problem

Android application developers face a huge manual burden when reviewing mobile app crowdsourcing test reports, resulting in inefficiency and possible omissions or misjudgments.

Method used

By automating the feature analysis of text information and screenshot images in the test report, repetitive reports are identified using similarity metrics between reports, and sorting them based on the test report's ability to reveal new defects, reducing the burden of manual review.

Benefits of technology

It significantly improves the speed and adequacy of application developers to obtain defective information, improves the efficiency of crowdsourcing testing of mobile applications, and reduces the error rate of manual review.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114780373B_ABST
    Figure CN114780373B_ABST
Patent Text Reader

Abstract

A method for sorting crowdsourced test reports for mobile applications based on image-text fusion analysis is characterized in that the method automatically extracts image and text features of crowdsourced test reports and sorts crowdsourced test reports according to the similarity measurement between reports, providing a new solution to the problem of excessive burden of manual review of test reports. The extracted image and text features will be recombined with defect-type features and context-type features to calculate similarity respectively. Defect similarity consists of problem control image similarity and defect description similarity, which is used to represent the information directly related to the defect displayed in the report. Context similarity consists of reproduction step similarity and context control similarity, which represents context information, including operation tracking that triggers the defect and activity information when the defect occurs. Finally, duplicate reports will be identified based on the similarity between test reports and sorted according to the ability of the reports to reveal new defects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of software testing, in particular to the field of Android crowdsourcing testing. Crowdsourcing workers perform manual testing on the required application to be tested and provide a test report including defect screenshots and defect descriptions. Background Art

[0002] With the development of mobile Internet and the development of a large number of Android applications, the general public has higher requirements for the quality of application software. The rapid iteration and frequent changes of mobile application software have brought many difficulties to traditional software testing methods. Testers often need to test each version of the application separately and quickly carry out multiple large-volume testing tasks in a short application development cycle, resulting in low testing efficiency and other problems. In view of the limitations of traditional mobile application software testing methods, testers have widely used mobile application crowdsourcing testing methods in many application scenarios.

[0003] Mobile application crowdsourcing testing requires testers to write reports containing defect screenshots and defect descriptions as results. However, the review efficiency of crowdsourcing test reports is a serious problem. There are a large number of duplicate reports in the test reports submitted by crowdsourcing workers. It will be a considerable burden for application developers to manually check a large number of test reports and remove duplicate reports to filter out valid information. Therefore, the present invention plans to perform a combined feature analysis on the text information and screenshot images in the test reports, identify duplicate reports through similarity metrics between reports, and sort them according to the ability of test reports to reveal new defects, so as to better help application developers to review a large number of test reports and obtain defect information more quickly and fully. Summary of the invention

[0004] The problem to be solved by the present invention is: the problem of the manual burden of Android application developers reviewing crowdsourced test reports for mobile applications. Crowdsourced test reports will be the result of a large number of test reports submitted by all testers, including defect descriptions and defect screenshots. If all of them are manually reviewed by application developers, the workload will be very large, which will not only damage the efficiency of crowdsourcing testing, but also make it difficult to avoid omissions and misjudgments in the process of manual review. Our invention automatically performs feature analysis on text information and screenshot images in test reports, uses similarity metrics between reports to identify duplicate reports, and sorts test reports according to their ability to reveal new defects. Reviewing the reports in the sorted order will greatly improve the speed and adequacy of application developers in obtaining defect information, and effectively improve the efficiency of crowdsourcing testing of mobile applications.

[0005] The technical solution of the present invention is: a method for sorting crowdsourced test reports of mobile applications based on image-text fusion analysis, which is characterized in that the image and text features of the crowdsourced test reports are automatically extracted, and the crowdsourced test reports are sorted according to the similarity measurement between reports to provide a new solution to the problem of excessive burden of manual review of test reports. Test report text feature extraction classifies the text content into two types of information: reproduction steps and defect descriptions according to natural language processing technology. The reproduction steps will be further analyzed and converted into an "operation-object" sequence, and the defect description will be further tagged with part of speech to extract the problem control information. The test report similarity measurement will include two parts of similarity, namely defect similarity and context similarity. Defect similarity consists of the similarity of the problem control image and the similarity of the defect description, which is used to represent the information directly related to the defect displayed in the report. Context similarity consists of the similarity of the reproduction steps and the similarity of the context control, which represents context information, including the operation tracking that triggers the defect and the activity information when the defect occurs. Test report sorting is to identify duplicate reports based on the similarity between test reports, and sort them according to the ability of the reports to reveal new defects. The method is divided into the following steps:

[0006] 1) Test report text feature extraction, mainly extracting two types of information, including reproduction steps and defect description. The reproduction steps will be further analyzed to extract the "operation-object" sequence, and the defect description will be further analyzed to extract the problem control description;

[0007] 1.1) Text content classification: Use the jieba Chinese word segmentation tool to perform Chinese word segmentation on the test report text content; remove stop words from the word segmentation results according to the pre-set stop word list; use Word2Vec to convert the word segmentation results into 128-dimensional numerical vectors; use the pre-trained TextCNN deep learning model to divide the text content into two categories: reproduction steps and defect description;

[0008] 1.2) Problem control description identification: Use the HMM-based text segmentation algorithm to segment the defect description; perform lexical analysis on each segmented text segment to extract the target noun field as the problem control description;

[0009] 1.3) Extraction of “operation-object” sequence: Also use the HMM-based text segmentation algorithm to segment the recurrence steps; perform lexical analysis on each segmented text segment; collect the verbs and corresponding objects in all segments and splice them into an “operation-object” sequence;

[0010] 2) Test report image feature extraction, mainly extracting two types of information, including problem control images and context controls;

[0011] 2.1) Problem control image extraction: Based on the problem control description information obtained in step 1.2, if a control in the screenshot contains the text content, use OCR technology to match the control with the problem control description and identify the control as the problem control image; if all controls in the screenshot do not contain the corresponding text content, use the deep learning model to analyze the intent of the control, match the intent text with the problem control description, and find the problem control image;

[0012] 2.2) Contextual control image extraction: In addition to the problematic control image obtained in step 2.1, the remaining controls analyzed in the screenshot are classified as contextual controls. The type of each contextual control is analyzed using the pre-trained CNN deep learning model. There are 14 different types of controls in total. Therefore, a 14-dimensional numerical vector will be obtained by analyzing all contextual control images. Each bit indicates how many contextual controls of this type are included in the screenshot image.

[0013] 3) Test report feature aggregation: All the features obtained in step 1 and step 2 are merged into two types of features: one is defect features, which refers to the characteristics that directly reflect or describe the defects in the crowdsourced test report; the other is context features, which consists of features that can provide a description of the environment when the defect occurs.

[0014] 3.1) The defect feature consists of the defect description obtained in step 1.1 and the image of the problem control obtained in step 2.1;

[0015] 3.2) The context feature is composed of the "operation-object sequence" obtained in step 1.3 and the context control image obtained in step 2.2;

[0016] 4) Test report similarity: Test report similarity is calculated by weighting defect similarity and context feature similarity, that is, test report similarity = γ*defect similarity + (1-γ)*context similarity;

[0017] 4.1) Defect similarity calculation: Defect similarity calculation consists of two weighted parts: defect description similarity and problem control image similarity. Problem control image similarity is calculated by extracting feature point sets using the SIFT algorithm. The defect description obtained in step 1.1 is converted into a numerical vector using Word2Vec, measured using Euclidean distance, and standardized. Defect similarity = α*problem control image similarity + (1-α)*defect description similarity.

[0018] 4.2) Calculation of context similarity: Context similarity is composed of the weighted similarity of the recurrence steps and the context control similarity; the context control similarity calculates the Euclidean distance of the context control type numerical vector obtained in step 2.2 and standardizes it; the recurrence steps similarity calculates the similarity of the "operation-object" sequence obtained in step 1.3 through the DTW algorithm and standardizes it; context similarity = β*context control similarity + (1-β)*recurrence steps similarity;

[0019] 5) Test report sorting: Calculate the report similarity dissimilarity matrix according to step 4, identify duplicate reports and sort the reports according to their ability to reveal new defects;

[0020] 5.1) Create a blank report; calculate the similarity between all test reports and the blank report according to the similarity calculation method in step 4, and select the test report with the lowest similarity to the blank report as the first report in the sorting sequence;

[0021] 5.2) Compare the average similarity of all reports with all reports in the sorted sequence, and select the report with the lowest average similarity to insert into the sorted sequence;

[0022] 5.3) Repeat step 5.2 until all reports have been successfully sorted;

[0023] The present invention is characterized in that:

[0024] 1. We propose a novel approach to prioritize crowdsourced test reports through deep screenshot understanding and detailed text analysis;

[0025] 2. We built a large-scale dataset group to deeply understand screenshots, including a large-scale control image dataset, a large-scale test report keyword vocabulary, a large-scale text classification dataset, and a large-scale crowdsourced test report dataset;

[0026] 3. Automatically extract test report features and calculate similarity, identify duplicate test reports based on this, and sort reports based on the test report's ability to reveal new defects;

[0027] Based on the above three points, the present invention can effectively solve the problem of excessive burden of manual review of mobile application crowdsourcing test reports, significantly improve the speed and adequacy of application developers in obtaining defect information, thereby improving the efficiency and effectiveness of mobile application crowdsourcing testing. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 The overall structure diagram of the present invention is DETAILED DESCRIPTION

[0029] The technologies involved in the present invention include jieba, Word2Vec, TextCNN, OCR, CNN, SIFT and DTW.

[0030] 1.jieba

[0031] Jieba is a Chinese word segmentation component of Python. This paper mainly uses the Chinese word segmentation and part-of-speech tagging functions of the tool for the conversion of numerical vectors and the identification of relevant feature information.

[0032] 2. Word2Vec

[0033] Word2vec is a toolkit for obtaining word vectors launched by Google in 2013. It is simple and efficient, and provides two language models: Skip-gram and CBOW, which can convert input words into numerical vectors.

[0034] 3. TextCNN

[0035] The TextCNN model is a model proposed by Yoon Kim in the paper Convolutional Naural Networks for Sentence Classification that uses convolutional neural networks to handle NLP problems. Compared with traditional RNN / LSTM models, TextCNN can more efficiently extract important features, which play an important role in classification.

[0036] 4. OCR

[0037] OCR stands for Optical Character Recognition, which analyzes and recognizes image data containing text content to obtain the text in it. Today's OCR technology is widely used and is mainly divided into general OCR technology and special OCR technology. General OCR technology handles more complex scenarios, while special OCR technology only recognizes a specific type of image data.

[0038] 5. CNN

[0039] Convolutional neural network is a type of feedforward neural network with deep structure and convolutional calculation, and is one of the representative algorithms of deep learning. Convolutional neural network has the ability of representation learning and can classify input information in a translation-invariant manner according to its hierarchical structure, so it is also called "translation-invariant artificial neural network".

[0040] 6. SIFT

[0041] SIFT stands for Scale Invariant Feature Transform. This description method in the field of image processing can detect a set of key points in an image. These key points are invariant in scale and are a local feature description operator. The SIFT algorithm has many advantages and can show good stability and invariance for image operations such as rotation, scaling, and brightness changes.

[0042] 7. DTW

[0043] The DTW (Dynamic Time Warping) algorithm is based on the idea of ​​dynamic programming and solves the template matching problem of different pronunciation lengths. It is an earlier and more classic algorithm in speech recognition and is used for isolated word recognition. The present invention mainly uses the DTW algorithm to calculate the similarity between different "operation-object" sequences.

[0044] In order to verify the practical usability of the present invention, a verification experiment was designed to analyze the effectiveness of the test report sorting results. 536 crowdsourced test reports covering 10 Android applications were selected for the experiment. The details of the applications are shown in Table 1. The present invention was used to sort all the test reports of each application. APFD (average percentage of fault detection) was used as the metric. In the formula, T fi represents the index of the first report in which defect i was found, n is the total number of reports, and M is the total number of defects found.

[0045]

[0046] In order to better demonstrate the advantages of the present invention compared with other methods, we set up several control groups, including:

[0047] (1) Ideal: This strategy is the best sorting method in theory, which means that developers can check all defects shown in the report in the shortest time.

[0048] (2) Image: This strategy only uses the deep image understanding results in our paper to rank the test reports, because deep image understanding is an important part of our research.

[0049] (3) Random: This strategy refers to the result of using random sorting without any priority strategy.

[0050] Table 1 Mobile application crowdsourcing testing Android application information table

[0051]

[0052] The experimental results of each application are shown in Table 2. For Ideal, we manually calculated the optimal situation of each cluster using the APFD formula. For the Random strategy, we took the average of 100 running results to eliminate randomness. For the present invention and the Image strategy used alone, we only need to run it once because of the stability of our method. Compared with the Random strategy, the optimization range of the present invention is from 15.15% to 38.93%, and the average improvement reaches 27.04%. Then, we compare the results of the present invention with those of the Image strategy used alone. The average improvement of the present invention is 4.54%, and in two applications (A3 and A4), the ranking results of the present invention are much better. For A8, the present invention is slightly weaker than the Image strategy. We checked the report of A8 and found that the quality of the text description written by the tester was not high and could not actively help the priority sorting of the report. In general, the results prove the necessity of combining text analysis and deep image understanding, which can make up for each other's shortcomings and improve the accuracy of priority sorting.

[0053] Table 2 Results of application defect verification experiments

[0054] serial number Ideal The present invention Image Random A1 0.974 0.927 0.926 0.805 A2 0.931 0.839 0.805 0.655 A3 0.957 0.865 0.721 0.692 A4 0.980 0.933 0.845 0.794 A5 0.967 0.898 0.892 0.751 A6 0.942 0.827 0.827 0.619 A7 1.000 1.000 1.000 0.750 A8 0.941 0.850 0.858 0.694 A9 0.958 0.931 0.847 0.681 A10 0.938 0.863 0.854 0.621

[0055] The present invention analyzes the actual report sorting results through specific experiments, verifies the effectiveness of the feature extraction results of the mobile application crowdsourcing test report and helps application developers to find all defects in the test report as early as possible. Application developers can review all reports in the order sorted by the present invention to improve the speed of finding defects and the adequacy of defect information.

Claims

1. A method for sorting mobile application crowdsourcing test reports based on image-text fusion analysis, characterized in that By automatically extracting the image and text features of crowdsourced test reports, the crowdsourced test reports are sorted according to the similarity measurement between the reports, and a priority sorting sequence of the test reports is generated; the steps of the report sorting method are as follows: 1) Test report text feature extraction, mainly extracting information including reproduction steps and defect description. The reproduction steps will be further analyzed to extract the "operation-object" sequence, and the defect description will be further analyzed to extract the problem control description. The specific steps include the following: 1.1) Text content classification: Segment the test report text content, remove stop words, and vectorize it, and then divide it into two categories: reproduction steps and defect description; 1.2) Problem control description identification: segment the defect description, and perform lexical analysis on each segmented text segment to extract the target noun field as the problem control description; 1.3) Extraction of "operation-object" sequence: segment the recurrence steps, perform lexical analysis on each segmented text segment, collect the verbs and corresponding objects in all segments and splice them into "operation-object" sequence; 2) Test report image feature extraction, mainly extracting information including problem control images and context controls, mainly including the following steps: 2.1) Problem control image extraction: Based on the problem control description information, if a control in the screenshot contains the text content, the control is matched with the problem control description and the control is identified as the problem control image; if all controls in the screenshot do not contain the corresponding text content, the intent of the control is analyzed, the intent text is matched with the problem control description, and the problem control image is found; 2.2) Contextual control image extraction: In addition to the problem control image, the remaining controls analyzed in the screenshot image are classified as contextual controls, the control type of the contextual control is identified, and each type of control is counted to form a contextual control type numerical vector; 3) Test report feature aggregation: All the features obtained in step 1 and step 2 are merged into two types of features: one is defect features, which refers to the characteristics that directly reflect or describe the defects in the crowdsourced test report; the other is context features, which consists of features that can provide a description of the environment when the defect occurs. The specific explanations are as follows: 3.1) The defect feature consists of the defect description obtained in step 1.1 and the image of the problem control obtained in step 2.1; 3.2) The context feature is composed of the "operation-object sequence" obtained in step 1.3 and the context control image obtained in step 2.2; 4) Test report similarity calculation: Test report similarity is calculated by weighted calculation of defect similarity and context feature similarity, that is, test report similarity = γ*defect similarity + (1-γ)*context similarity; 4.1) Defect similarity calculation: Defect similarity calculation is composed of the weighted combination of defect description similarity and problem control image similarity; Defect similarity = α*problem control image similarity + (1-α)*defect description similarity; 4.2) Calculation of context similarity: Context similarity is composed of the weighted combination of the similarity of the recurrence steps and the similarity of the context controls; context similarity = β * context control similarity + (1-β) * recurrence steps similarity; 5) Test report sorting: Calculate the report similarity dissimilarity matrix according to step 4, identify duplicate reports and sort the reports according to their ability to reveal new defects; 5.1) Create a blank report; calculate the similarity between all test reports and the blank report according to the similarity calculation method in step 4, and select the test report with the lowest similarity to the blank report as the first report in the sorting sequence; 5.2) Compare the average similarity of all reports with all reports in the sorted sequence, and select the report with the lowest average similarity to insert into the sorted sequence; 5.3) Repeat step 5.2 until all reports have been successfully sorted.

2. The method for sorting mobile application crowdsourcing test reports based on image-text fusion analysis according to claim 1 is characterized in that: 1) In text content classification, the word segmentation tool used is Jieba, the word vector is constructed using Word2Vec, each word vector is 128 dimensions, and the classification uses the pre-trained TextCNN model; 2) In the problem control description recognition and "operation-object" sequence extraction steps, the text segmentation method adopts the HMM-based text segmentation algorithm; 3) In the problem control image extraction step, the text in the image is extracted using OCR technology; 4) In the context control image extraction step, the control type is implemented using a pre-trained CNN-based control classification model; 5) In the defect similarity calculation process, the similarity of the problem control image is calculated by extracting the feature point set through the SIFT algorithm, and the similarity of the defect description text is measured by the Euclidean distance of the word vector and standardized; 6) In the context similarity calculation, the context control similarity is defined by the Euclidean distance between the context control type numeric vectors and is standardized; the repetition step similarity is calculated by the DTW algorithm "operation-object" sequence similarity and is standardized.

Citation Information

Patent Citations

  • Image-based computer identification device and method for mobile crowdsourcing test report

    CN110363248A

  • Crowdsourcing test report processing method and device

    CN113220565A