A traceable function view construction method for crowd-sourced test reports

CN122673044APending Publication Date: 2026-09-01NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610759377.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-23
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

[0006]本发明的目的在于提供一种面向众包测试报告的可追溯功能视图构建方法、装置、电子设备及存储介质,以解决现有技术中以互斥聚类为核心的众包测试报告组织方式难以表达跨报告重叠证据、难以恢复跨簇关联且不利于人工审阅的问题

Benefits of technology

[0012] (1) The present invention changes the organization method of crowdsourced test reports from mutually exclusive clustering to the construction of overlapping functional views, so that the same report can support multiple functional views at the same time, which is more in line with the actual characteristics of cross-cluster distribution of fault evidence in crowdsourced test reports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122673044A_ABST
    Figure CN122673044A_ABST
Patent Text Reader

Abstract

This invention discloses a method, apparatus, electronic device, and storage medium for constructing traceable functional views for crowdsourced test reports. The method includes: acquiring multiple crowdsourced test reports; performing evidence-anchored structured extraction on each report according to a preset entity pattern to obtain entity, relationship, and source evidence mappings, and constructing an original evidence graph; performing unified processing on entities of the same type in the original evidence graph under entity type constraints to obtain a unified evidence graph; selecting preset seed entities from the unified evidence graph, performing local expansion, path filtering, candidate assembly, and candidate deduplication to generate multiple candidate functional views; and selecting and outputting a representative set of functional views under budget constraints based on the clue coverage and review cost of the candidate functional views. This invention enables cross-report, overlapping fault evidence organization while preserving the traceability of report sources, improving the efficiency of defect triage, cross-cluster association recovery, and manual review of crowdsourced test reports.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of software testing, defect analysis, natural language processing, and graph data processing, and particularly to a method, apparatus, electronic device, and storage medium for constructing a traceable functional view for crowdsourced test reports. Background Technology

[0002] Crowdsourced testing can improve software defect discovery by introducing different testers, equipment, and execution environments, but it also generates a large number of heterogeneous and noisy natural language test reports. Developers face a bottleneck in practical applications when triaging these reports, as they are subjected to a large number of reports with inconsistent styles, completeness, terminology, and irrelevant descriptions, and extracting actionable fault evidence from them.

[0003] Existing solutions typically employ a report-centric approach, such as classifying, sorting, clustering, summarizing, or selecting representative items based on report-level features, report-level similarity, or report-level grouping. The process often follows a "divide first, then analyze" approach, that is, first divide the reports into mutually exclusive clusters, and then perform subsequent tasks around a single report or cluster.

[0004] However, in crowdsourced testing scenarios, evidence related to the same functional failure is often scattered across reports and may cross cluster boundaries, while different failures may be incorrectly aggregated due to similar surface clues. Report organization methods based on mutually exclusive clustering cannot accurately reflect the structural characteristics of evidence intersection, overlap, and cross-cluster distribution in crowdsourced testing reports.

[0005] Therefore, there is an urgent need for a crowdsourced test report processing solution that can maintain source auditability, support cross-report correlation recovery, and adapt to limited review budgets. Summary of the Invention

[0006] The purpose of this invention is to provide a method, apparatus, electronic device, and storage medium for constructing a traceable functional view for crowdsourced test reports, in order to solve the problems in the existing crowdsourced test report organization method based on mutually exclusive clustering that is difficult to express cross-report overlapping evidence, difficult to restore cross-cluster associations, and not conducive to manual review.

[0007] To achieve the above objectives, this invention provides a method for constructing traceable functional views for crowdsourced test reports. The core idea is to transform the original crowdsourced test report into an evidence graph and form a set of representative functional views under budget constraints through four stages: "extraction-unification-guidance-selection".

[0008] In a preferred embodiment, each crowdsourced test report is first subjected to evidence-anchored structured extraction to construct an original evidence graph consisting of entity, relation, and source mapping. Each retained entity is bound to an identifiable text fragment in the original text, and each retained relation is directly supported by the text or connected to an entity with traceable source.

[0009] In another preferred embodiment, the entities in the original evidence graph are subject to type constraints to unify the graph, allowing unification only between entities of the same type, and equivalence verification is performed based on anchor references, source contexts, and type-consistent neighborhoods to form a unified evidence graph.

[0010] In another preferred embodiment, the problem statement entity, the observed symptom entity, and the system diagnosis entity are used as seed entities. On the unified evidence graph, local expansion, path filtering, candidate assembly, and candidate deduplication are performed to construct multiple candidate functional views. Each candidate functional view is a locally connected subgraph associated with the fault clues and retains the mapping between the supporting report set and the text evidence fragment set.

[0011] In another preferred embodiment, the candidate functional views are subjected to semantic redundancy compression and representative selection under budget constraints. Based on a comprehensive score of new clue coverage and review cost, a set of representative functional views within a given budget is iteratively selected.

[0012] (1) The present invention changes the organization method of crowdsourced test reports from mutually exclusive clustering to the construction of overlapping functional views, so that the same report can support multiple functional views at the same time, which is more in line with the actual characteristics of cross-cluster distribution of fault evidence in crowdsourced test reports.

[0013] (2) This invention maintains the traceability of nodes, edges and final functional views to source reports and source text fragments through evidence anchoring extraction and source mapping, which facilitates manual auditing and result backtracking.

[0014] (3) This invention improves the fragmentation of the original graph and enhances the graph connectivity and coverage potential by unifying the entities through type constraints.

[0015] (4) The present invention introduces a budget constraint and cost normalization mechanism in view selection, which can retain more manageable and reviewable functional perspectives under limited review resources, and avoid view collapse caused by a few large views occupying the entire budget.

[0016] (5) The functional view formed by the present invention can also support cross-cluster retrieval and association recovery under fixed cluster boundaries. Attached Figure Description

[0017] Figure 1This is a schematic diagram comparing the existing mutually exclusive clustering organization method with the functional view organization method of the present invention.

[0018] Figure 2 This is a flowchart of the overall processing in an embodiment of the present invention, showing the four stages of "extraction, unification, induction, and selection". Detailed Implementation

[0019] Example 1: Overall Method Flow

[0020] The input in this embodiment is multiple test reports generated by a crowdsourced testing platform. Each test report can be represented as a tuple:

[0021] r_i = (id_i, x_i)

[0022] Here, id_i represents the report identifier, and x_i represents the natural language text content of the report. The goal of this embodiment is not to divide all reports into mutually exclusive clusters, but rather to organize entities related to fault clues and their supporting evidence into a set of functional views S* under budget constraints. Each functional view is a representative subgraph and preserves the mapping relationship between supporting reports and source text fragments.

[0023] In this embodiment, the overall processing flow includes four stages: the first stage is evidence-anchored structured extraction; the second stage is type-constrained entity unification; the third stage is candidate functional view induction; and the fourth stage is representative functional view selection under budget constraints.

[0024] Example 2: Evidence-Anchored Structured Extraction

[0025] In this embodiment, the crowdsourcing test report is first extracted in a structured manner according to a preset entity pattern. Preferably, the preset entity pattern includes the following four core evidence types: 1. FunctionalModule entity, denoted as MOD; 2. UserAction entity, denoted as OP; 3. ObservedSymptom entity, denoted as PHEN; 4. SystemDiagnosis entity, denoted as DIAG.

[0026] The above four types of entities constitute the chain of evidence:

[0027] MOD->OP->PHEN->DIAG

[0028] In a further implementation, the following two types of extended entities may be introduced: 1. ImpactElement, denoted as ELEM; 2. ProblemStatement, denoted as ISSUE. These extended entities are used for view naming, organization, and coverage statistics. When there is no explicit mention in the original text, it can be derived from sufficiently anchored core evidence through a deterministic source preservation approach.

[0029] This embodiment also limits the types of relationships that can be instantiated during the extraction phase, including: (1) context anchoring relationship, used to attach user operations or fault clues to functional modules; (2) interaction target relationship, used to connect OP and ELEM; (3) result chain relationship, used to connect local operation-symptom-diagnosis chain; (4) support / derivation relationship, used to associate analysis layer nodes such as ISSUE or ELEM with underlying core evidence.

[0030] In terms of formal implementation, a structured evidence record is generated for each report r_i:

[0031] X(r_i)=(V_i,E_i,H_i)

[0032] Where V_i is the extracted entity set, E_i is the extracted in-report relation set, and H_i is the source mapping, used to record the supporting text fragments corresponding to each node or edge; the evidence records from multiple reports are aggregated to obtain the original evidence graph:

[0033] G_raw=(V,E), V=∪V_i, E=∪E_i

[0034] And the global source mapping ∏=∪Π_i. Additionally, define the report support set for any graph object y:

[0035]

[0036] In a preferred embodiment, after the extraction stage follows the large language model or other structured extraction model, a conservatism check is performed through the extraction safeguard layer, including type validity check, relation compatibility check, source fragment recoverability check, and deduplication, thereby filtering out results that do not meet the evidence constraints.

[0037] Example 3: Type-Constrained Entity Unification

[0038] Since the original evidence graph is still a mention-level graph structure, semantically equivalent concepts in different reports may appear in different superficial forms, resulting in insufficient cross-report connectivity. Therefore, this embodiment applies type constraint entity unification to the original evidence graph to elevate the graph to a concept-level structure while preserving source traceability.

[0039] set up Let T be the set of entity nodes and T be the set of entity types. Then the unified normalized mapping of entities is:

[0040] g:E->E*

[0041] Where E* represents the normalized set of entity representatives. Applying the mapping g, we obtain the unified evidence graph:

[0042] G_unify = ApplyMao(G_raw, g)

[0043] The unified entity selection process includes four stages: semantic bucketing, candidate entity pair proposal, evidence-aware equivalence verification based on anchored mentions, source fragments, and neighborhood information, and closure-based merging by connected components.

[0044] In this embodiment, only entities of the same type are allowed to participate in the unification. For a given entity type t, we have:

[0045] E_t={e∈E|type(e=t}

[0046] First, the search space is narrowed down within E_t by recalling and guiding buckets, and then candidate pairs are rigorously verified. Candidate pairs with confidence scores below a preset threshold are not merged to avoid the spread of errors to subsequent stages. After unification, the source mappings of nodes and edges are retained in Graph A.

[0047] Example 4: Candidate Function View Guidance

[0048] While the unified evidence graph offers stronger cross-report connectivity, it remains too large to be a direct output for manual triage. Therefore, this embodiment further induces candidate functional views on the unified evidence graph.

[0049] Preferably, the following seed entity types are selected:

[0050] T_seed={ISSUE, PHEN, DIAG}

[0051] Because these nodes more directly carry fault clues, they are suitable as entry points for the functional view. For a seed node s, it can be considered a valid seed if its type belongs to the above set, it has a non-empty source mapping, and there is at least one type-compatible context node in its local neighborhood.

[0052] In this embodiment, the following steps are performed sequentially for each valid seed: (1) Local expansion: Bounded hop expansion is performed around the seed node on the unified graph; (2) Path filtering: Only nodes and edges that satisfy the locality constraint and can be connected to the interpretive context are retained; (3) Candidate assembly: The retained local structures are assembled into a connected subgraph S = (V_S, E_S); (4) Filtering and deduplication: Candidate views that are too small, have weak support, or have approximately duplicate evidence signatures are deleted.

[0053] Locality optimization is achieved through two types of constraints: a hop count cap to limit expansion drift and a report support cap to prevent overly generalizable clues from dominating the candidate set.

[0054] Simultaneously, define the source mapping and reporting support set for candidate views:

[0055] ∏(S)=∪∏(y),Γ(S)=∪Γ(y),y∈V_S∪E_S

[0056] And define the set of clue contributions for this candidate view:

[0057] C(S)={v∈V_S|type(v)∈T_seed}

[0058] The valid candidate feature view constitutes the candidate pool C = {S|Valid(S)}.

[0059] Example 5: Selection of Representative Functional Views under Budget Constraints

[0060] Candidate pools typically have high recall but also contain redundant results. To provide a representative set of views that can be reviewed with limited review resources, this embodiment performs view selection under budget constraints. Preferably, the candidate pool can be semantically merged first to compress nearly duplicated views before coverage optimization is performed.

[0061] Let the compressed candidate pool be C′, and the complete set of clues be U={u∈V|type(u)∈T_seed}. For any candidate view S∈C′, its contribution is... The objective under budget constraints can be expressed as:

[0062] max|∪C(S)|st∑cost(S)≤B

[0063] Where B represents the review budget. In a preferred embodiment, the review cost of a candidate view is defined as:

[0064] cost(S)=|Γ(S)|+λ·(|V_S|+|E_S|), λ≥0

[0065] Here, |Γ(S)| approximately represents the number of source reports that reviewers may need to look at, and |V_S|+|E_S| represents the structure size penalty, used to suppress candidate views that are too large and difficult to read, even though they are supported by the same batch of reports.

[0066] Since the above optimization problem is NP-hard, a greedy selector is preferred. Let Δ(S) be the marginal increase in cue coverage brought by the candidate view at the current stage, then the score can be calculated using the following formula:

[0067] score(S)=Δ(S) / (cost(S)+c_0)

[0068] Then, within the remaining budget, iteratively add the highest-rated feasible candidate views until the budget is exhausted or no new leads are available. The final output S* is a set of representative, overlapping functional views under budget constraints, rather than a mutually exclusive partition of the report.

[0069] Example 6: Method Effect Example

[0070] In the corresponding experiments, the graph structure of multiple application datasets was improved after entity unification, the number of connected components decreased, the proportion of the largest connected core increased, and the coverage capability under fixed top-K selection was improved.

[0071] Under a shared budget, a complete configuration scheme can retain more manageable views within the same budget, while removing cost normalization makes it easier for the candidate set to concentrate on a few large views.

[0072] Under fixed text clustering boundaries, the retrieval results based on the functional view can recover more cross-cluster related reports, indicating that the functional view formed by the present invention is not only applicable to report organization, but also to cross-cluster retrieval.

[0073] Further manual task audits show that, compared to cluster-based organization, the functional view organization of this invention is more practical in terms of completion time, topic identification, and supporting evidence location.

[0074] Example 7: Device, Electronic Equipment, and Storage Medium

[0075] The present invention also provides an apparatus for implementing the above method, which can be deployed in a crowdsourcing testing platform, a defect management platform, a software quality analysis system, or a server-side processing system. The apparatus includes at least: an acquisition module, an extraction and mapping module, an entity unification module, a view guidance module, a budget selection module, and an output module.

[0076] The module consists of: an acquisition module for receiving or reading crowdsourced test reports; an extraction and mapping module for performing evidence-anchored extraction and constructing the original evidence graph; an entity unification module for performing type-constrained entity unification and retaining source mapping; a view guidance module for generating candidate functional views around seed entities; a budget selection module for performing redundancy compression and representative view selection under budget constraints; and an output module for outputting representative functional views and their corresponding source reports, text fragments, and traceable relationships.

[0077] The present invention also provides an electronic device, including a processor, a memory, and a computer program stored in the memory, wherein when the processor executes the computer program, it implements the method described in any of the above embodiments.

[0078] The present invention also provides a computer-readable storage medium, wherein The system stores a computer program, which, when executed by a processor, implements the method described in any of the above embodiments.

[0079] This invention revolves around crowdsourced test reports. To address the practical problem of evidence being scattered across reports and clusters with overlapping capabilities, this paper proposes a processing scheme consisting of evidence anchoring extraction, unified type constraints, functional view guidance, and budget constraint selection. Instead of using mutually exclusive clustering as the sole output, this scheme constructs a traceable, overlapping, and reviewable set of functional views, making it more suitable for supporting defect triage, cross-cluster correlation recovery, and manual review.

[0080] Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this invention shall be included within the scope of protection of this invention.

Claims

1. A method for constructing a traceable functional view for crowdsourced test reports, characterized in that, include: S1. Obtain multiple crowdsourcing test reports, each of which must include at least a report identifier and natural language text content; S2. According to the preset entity pattern, perform evidence anchoring structured extraction on each crowdsourcing test report to obtain an entity set, a relation set, and a source evidence set corresponding to the entity set and relation set. Construct an original evidence graph based on the entity set, relation set, and source evidence set. The source evidence set records the mapping relationship between nodes or edges and source report identifiers and text fragments. S3. Perform type constraint unification processing on the entities in the original evidence graph to obtain a unified evidence graph. The type constraint unification processing includes at least: candidate recall binning of entities of the same type; generating candidate entity pairs; performing equivalence verification based on the anchor mentions, source evidence context, and neighborhood structure information corresponding to the candidate entity pairs; and verifying the candidate entity pairs that pass the verification. S4. Merge the unified evidence graph according to the connected components and retain the mapping relationship between the unified entity and the original source evidence; S5. Select a preset seed entity from the unified evidence graph, and perform local expansion, path filtering, candidate assembly and candidate deduplication based on the preset seed entity to generate multiple candidate functional views, wherein each candidate functional view is a connected subgraph associated with the fault clue and is associated with the source report set and text evidence fragment set supporting the candidate functional view; S6. Based on the clue coverage and review cost corresponding to each candidate functional view, select a representative functional view set from multiple candidate functional views under budget constraints; S7. Output the representative functional view set, as well as the source report identifier set and text evidence fragment set corresponding to each representative functional view.

2. The method according to claim 1, characterized in that, The preset entity pattern includes at least one or more of the following core entity types: functional module entity, user operation entity, observed symptom entity, and system diagnosis entity.

3. The method according to claim 2, characterized in that, The preset entity pattern also includes one or more of the following extended entity types: influence element entities and question statement entities; wherein the extended entities are obtained through explicit textual reference extraction or derived from anchored core evidence in a deterministic manner that maintains traceability of origin.

4. The method according to claim 1, characterized in that, The set of relationships obtained in step S2 includes at least one of the following: (1) context anchoring relationships that attach user actions or fault clues to functional modules; (2) interactive target relationships between user actions and influencing elements; (3) result chain relationships that connect user actions, observed symptoms and system diagnoses; and (4) support or derivation relationships that associate extended entities with core evidence.

5. The method according to claim 1, characterized in that, In step S3, uniform processing is only allowed between entities of the same type; when the equivalence verification result of the candidate entity pair is lower than the preset threshold, merging is not performed.

6. The method according to claim 1, characterized in that, The preset seed entity mentioned in step S4 is selected from at least one or more of the problem statement entity, the observed symptom entity, and the system diagnosis entity.

7. The method according to claim 1, characterized in that, The review cost in step S5 is determined by the number of source reports associated with the candidate functional view and the structural size of the candidate functional view, wherein the structural size includes at least one of the number of nodes and the number of edges.

8. The method according to claim 1, characterized in that, Different functional views within the representative functional view set are allowed to share some entities, some evidence, or some source reports, so that the same crowdsourced test report can support multiple functional views simultaneously.

9. An electronic device comprising a processor and a memory, wherein the memory stores a computer program, which, when executed by the processor, implements the method of any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method of any one of claims 1 to 8.