Mock data life cycle governance and replayable method and system
Patent Information
- Application Number
- CN202610459171.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-09
- Publication Date
- 2026-09-29
AI Technical Summary
[0008]本发明实施例中提供一种Mock数据生命周期治理与可重放方法,以解决现有技术中Mock数据管理方式通常存在静态Mock文件易老化,维护成本高;直接访问真实环境导致不确定性,难以可重放;缺乏以“契约”为核心的版本化治理与差异分析;缺少闭环演进与审计追溯的问题
本发明提供了一种Mock数据生命周期治理与可重放方法及系统,其中,该方法通过在测试执行时将目标快照注入作为响应,隔离真实环境数据动态性与外部依赖波动,使相同测试用例在不同时间、不同环境中能够获得一致输入输出,从而实现稳定可重放,提升回归测试的稳定性与可复现性。通过对命中白名单的接口进行持续采集,形成多个时间点的Mock快照,并以契约摘要哈希值与采集时间戳构建双级索引存储,使Mock数据从一次性样例转化为可治理资产。新快照产生时还可执行分层差异比对并标记差异性质,及时发现结构层变化与关键字段变化,避免偏差长期积累。通过对Mock快照提取骨架结构并计算契约摘要哈希值,能够以“结构签名”的方式聚类与识别结构变化;结合测试用例中间表示IR进行目标快照检索与契约一致性校验,并在失败时触发降级策略,使测试在接口变更过程中具有更强的韧性,减少大量人工维护与临时修复。通过稳定差异识别与JSON Patch补丁生成,仅对经噪声过滤并稳定出现的差异进行最小化修改,生成新的时间版本快照而非覆盖旧快照,有利于控制变更范围,避免快照整体漂移导致不可控;同时保留历史版本以支持回溯与复现。通过记录审计日志,将快照演进形成证据链,便于质量追踪、问题定位、团队协作与合规审计。
Smart Images

Figure CN122838262A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of software testing and data governance technology, specifically to a method and system for Mock data lifecycle governance and replayability. Background Technology
[0002] With the widespread adoption of microservice architecture, distributed systems, and continuous integration / continuous delivery (CI / CD) engineering practices, business systems often rely heavily on HTTP / RPC interface interactions. To reduce dependence on the real environment and improve testing efficiency and stability during the development, integration, and regression testing phases, a common practice is to introduce a mock mechanism, which intercepts target interface calls during test execution and returns pre-prepared response data.
[0003] However, existing Mock data management methods typically have the following shortcomings: 1) Static mock files are prone to aging and have high maintenance costs. Traditional solutions often store mock data in the form of JSON files, recording scripts, or fixed response samples. While this data may be sufficient for testing at the time of recording, as interface fields are added, deleted, types change, and enumeration values expand, static mocks gradually deviate from the actual contract, leading to false positives, missed positives, or outright failures in test cases. This necessitates manually updating a large number of mock files, and the lack of a unified standard in the update process easily introduces new inconsistencies and maintenance burdens.
[0004] 2) Direct access to the real environment introduces uncertainty and makes replayability difficult. Another approach involves directly requesting the test environment, pre-production environment, or even the production environment during testing to obtain the "latest data." This approach can alleviate "aging," but real-world environment data is dynamic (such as order status, inventory, promotional activities, timestamps, random IDs, etc.) and may also be affected by concurrency, fluctuations in dependent services, and changes in external systems. This can lead to unstable test inputs, making it difficult to reliably reproduce a particular failure. Once regression testing fails, it is difficult to determine whether the cause is a code defect, data change, or environmental fluctuation, resulting in low efficiency in troubleshooting.
[0005] 3) Lack of version-based governance and difference analysis centered on "contracts" Most mock solutions have rudimentary data version management, distinguishing data only by filename, time, or directory. They lack a "contract summary" or similar mechanism for the interface return structure, making it impossible to automatically determine structural compatibility when the interface changes. They also struggle to conduct systematic difference comparisons and mark the nature of the differences. The lack of effective difference visualization and impact assessment leads to the long-term accumulation of discrepancies between mock data and real interfaces, ultimately forming "dirty data assets" that are difficult to manage.
[0006] 4) Lack of closed-loop evolution and audit traceability When testing reveals mismatches in mock data, traditional solutions often rely on manual modification or replacement of the overall response data. This approach fails to guarantee minimal modifications and lacks comprehensive audit logs, making it difficult to trace "who, when, why, and which fields were modified." In a context where quality compliance and security audits are increasingly important, the lack of auditing poses significant risks.
[0007] Therefore, a governance mechanism is needed that can continuously collect Mock data, perform contractual modeling, store it in a two-level indexed versioned manner, retrieve and replay target snapshots based on IR, identify stable differences and evolve them through patching, and has audit logs, so as to achieve a closed loop of replayable and sustainable governance of Mock data. Summary of the Invention
[0008] This invention provides a method for Mock data lifecycle governance and replayability to address the problems of existing Mock data management methods, such as static Mock files being prone to aging and having high maintenance costs; direct access to the real environment leading to uncertainty and difficulty in replayability; lack of version governance and difference analysis centered on "contracts"; and lack of closed-loop evolution and audit traceability.
[0009] To achieve the above objectives, on the one hand, the present invention provides a method and system for Mock data lifecycle governance and replayability, wherein the method includes: S1, listening to requests of the current interface that hit a preset whitelist, collecting response data of each request, and performing desensitization processing on each response data to generate multiple Mock snapshots corresponding to the current interface; S2, parsing and extracting the skeleton structure of each Mock snapshot, and calculating the unique contract digest hash value corresponding to each Mock snapshot by using a hash algorithm for each extracted skeleton structure; S3. Using the contract digest hash value as the first-level index and the collection timestamp as the second-level index, store each Mock snapshot containing two levels of indexes in the snapshot database. S4. Before executing tests using test cases, obtain the intermediate representation (IR) of the test cases, and retrieve the target snapshot from the snapshot database corresponding to the current interface based on the intermediate representation IR and the two-level indexes. When a change to the current interface is detected, retrieve the target snapshot again according to the changed current interface. S5. During test execution, use the target snapshot as the response, and simultaneously acquire the real response for difference comparison and output a difference view. Identify stable differences between the real response and the target snapshot, generate patches for the identified stable differences, and apply the patches to the target snapshot to generate a new time-version Mock snapshot.
[0010] Optionally, S1 includes: monitoring network requests of the current interface, combining the requested URL path and HTTP method as a group, and performing regular expression matching with a preset whitelist; if the match is successful, continuously collecting response data from the current interface at preset intervals; if the match is unsuccessful, allowing the request to proceed without data collection; performing desensitization processing on sensitive fields in each collected response data to obtain multiple desensitized data; encapsulating each desensitized data with the corresponding collected metadata to form multiple Mock snapshots corresponding to the current interface; wherein, the collected metadata includes: a collection timestamp.
[0011] Optionally, the step of calculating the unique contract digest hash value corresponding to each Mock snapshot by hashing each extracted skeleton structure includes: sorting all field names in lexicographical order for each extracted skeleton structure, and generating a normalized string corresponding to each skeleton structure based on the sorted structure; calculating the unique contract digest hash value corresponding to each Mock snapshot by hashing each normalized string using a hash algorithm; and injecting a business tag into each Mock snapshot.
[0012] Optionally, after S3, the method further includes: whenever a new Mock snapshot is generated, performing a hierarchical difference comparison between the new Mock snapshot and the old Mock snapshot of the previous time version of the current interface, wherein the hierarchical difference comparison includes at least a structural layer difference comparison and a key field value layer difference comparison; outputting a difference report based on the comparison results and marking the nature of the differences.
[0013] Optionally, retrieving the target snapshot in the snapshot database corresponding to the current interface based on the intermediate representation IR and utilizing the two-level index includes: retrieving multiple candidate snapshots in the snapshot database corresponding to the current interface based on the intermediate representation IR and utilizing the two-level index; calculating the matching score between each candidate snapshot and the test case, and selecting the candidate snapshot with the highest score as the target snapshot; performing contract consistency verification on the target snapshot; and triggering a degradation strategy to select an alternative snapshot as the target snapshot or output an exception when the verification fails.
[0014] Optionally, the matching score is obtained by weighting at least the contract compatibility score, the tag matching score, and the time freshness score.
[0015] Optionally, the stable difference identification of the difference between the real response and the target snapshot includes: performing noise filtering on the difference between the real response and the target snapshot to exclude preset non-business field differences; statistically analyzing the remaining differences after filtering, and determining that the same difference is a stable difference when it appears in N consecutive test executions of the same interface.
[0016] On the other hand, this invention provides a Mock data lifecycle governance and replayability system, which includes: a snapshot generation unit, used to listen to requests of the current interface that hit a preset whitelist, collect response data of each request, and perform desensitization processing on each response data to generate multiple Mock snapshots corresponding to the current interface; a hash calculation unit, used to parse and extract the skeleton structure of each Mock snapshot, and calculate the unique contract digest hash value corresponding to each Mock snapshot by using a hash algorithm for each extracted skeleton structure; and a snapshot storage unit, used to store each Mock snapshot containing two levels of indexes, with the contract digest hash value as the first-level index and the collection timestamp as the second-level index. The system stores the target snapshot in a snapshot database. A target snapshot retrieval unit is used to obtain the intermediate representation (IR) of the test case before executing the test, and retrieve the target snapshot from the snapshot database corresponding to the current interface based on the IR and the two-level index. When a change to the current interface is detected, the target snapshot is retrieved again according to the changed interface. An evolution and auditing unit is used to use the target snapshot as a response during test execution, and to obtain the real response in parallel for difference comparison and output a difference view. Stable difference identification is performed on the difference between the real response and the target snapshot, and patches are generated for the identified stable differences. These patches are then applied to the target snapshot to generate a new time-version Mock snapshot.
[0017] Optionally, the snapshot generation unit includes: a matching subunit, used to listen to network requests of the current interface, and perform regular expression matching on the requested URL path and HTTP method as a combination with a preset whitelist; if the match is successful, the response data of the current interface is continuously collected at preset intervals; if the match is unsuccessful, the request is allowed directly without data collection; a desensitization subunit, used to perform desensitization processing on sensitive fields in each collected response data to obtain multiple desensitized data; and an encapsulation subunit, used to encapsulate each desensitized data together with the corresponding collected metadata to form multiple Mock snapshots corresponding to the current interface; wherein, the collected metadata includes: a collection timestamp.
[0018] Optionally, the step of calculating the unique contract digest hash value corresponding to each Mock snapshot by hashing each extracted skeleton structure includes: sorting all field names in each extracted skeleton structure lexicographically, and generating a normalized string corresponding to each skeleton structure based on the sorted structure; calculating the unique contract digest hash value corresponding to each Mock snapshot by hashing each normalized string using a hash algorithm; and injecting a business tag into each Mock snapshot.
[0019] The beneficial effects of this invention are: This invention provides a method and system for Mock data lifecycle governance and replayability. The method injects target snapshots as responses during test execution, isolating the dynamics of real-world data and fluctuations in external dependencies. This ensures consistent input and output for the same test cases across different times and environments, achieving stable replayability and improving the stability and reproducibility of regression testing. By continuously collecting data from whitelisted interfaces, multiple Mock snapshots are generated at various time points. A two-level index is constructed using contract digest hash values and collection timestamps, transforming Mock data from one-off samples into manageable assets. When a new snapshot is generated, hierarchical difference comparisons are performed and the nature of the differences is marked, promptly identifying structural changes and key field changes to prevent long-term bias accumulation. By extracting the skeleton structure from the Mock snapshots and calculating the contract digest hash value, structural changes can be clustered and identified using a "structural signature" approach. Combined with the intermediate representation (IR) of test cases, target snapshot retrieval and contract consistency verification are performed, triggering a degradation strategy upon failure. This makes testing more resilient during interface changes, reducing the need for extensive manual maintenance and temporary fixes. By identifying stable differences and generating JSON patches, only minimal modifications are made to differences that have been filtered out of noise and remain stable. This generates new time-based snapshots instead of overwriting old ones, which helps control the scope of changes and avoids uncontrollable shifts caused by overall snapshot drift. Historical versions are also retained to support backtracking and reproduction. By recording audit logs, the snapshot evolution is organized into a chain of evidence, facilitating quality tracking, issue localization, team collaboration, and compliance auditing. Attached Figure Description
[0020] Figure 1 This is a flowchart of a method for Mock data lifecycle governance and replayability provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of a Mock data lifecycle governance and replayable system provided in an embodiment of the present invention. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0022] Figure 1 This is a flowchart of a method for Mock data lifecycle governance and replayability provided in an embodiment of the present invention, such as... Figure 1 As shown, the method includes: S1. Listen for requests to the current interface that hit the preset whitelist, collect the response data of each request, perform de-identification processing on each response data, and generate multiple Mock snapshots corresponding to the current interface. In an optional implementation, S1 includes: S11. Monitor network requests for the current interface, combine the requested URL path and HTTP method into a single string, and perform regular expression matching against a preset whitelist. If a match is found, continuously collect response data from the current interface at preset intervals. If a match is not found, allow the request directly without collecting any data. In one optional implementation, the data collection module is deployed on the application side (e.g., SDK interceptor, middleware) or the gateway side (e.g., API Gateway, service mesh Sidecar). The data collection module listens for network requests received by the current interface (any interface). An interface is a service endpoint defined by a URL path pattern and an HTTP method. The URL path pattern describes the acceptable resource path format for the interface (which may include parameter placeholders), and the HTTP method specifies the type of operation to be performed on the resource (e.g., GET, POST, etc.). A network request is an actual call initiated by the client, containing a specific URL (whose path part is an instantiation of the interface path pattern) and an HTTP method. It is through this information in the request that the system can determine the target interface corresponding to the request. Whenever a request occurs, the URL path and HTTP method of the request are extracted, and they are combined and matched against rules in a pre-defined whitelist using regular expressions. The pre-defined whitelist (or simply whitelist) may contain multiple rules (i.e., entries in the whitelist), and each rule includes at least one URL path regular expression and one HTTP method regular expression. If a match is found, the data collection process is initiated; if a match is not found, the request is allowed without any data collection operation.
[0023] In this embodiment, the current interface is the order details interface, its URL path pattern is / api / v1 / orders / {orderId}, and the expected HTTP method is GET.
[0024] The whitelist rule example is as follows: URL regular expression: ^ / api / v1 / orders / \d+ Method: GET When a GET request is received with the URL path / api / v1 / orders / 10001, a successful match is found, and data collection is initiated. If the request accesses other paths such as / api / v1 / coupons / 123 or uses other HTTP methods (such as POST), data collection will not be performed.
[0025] This mechanism offers two advantages: first, it avoids pollution caused by collecting irrelevant data; and second, it avoids performance degradation of the overall system.
[0026] Once a request is whitelisted, data is continuously collected from the current interface at preset intervals. For example, it can be configured to sample the response data of requests that match the whitelist at regular intervals; or to collect the response data of requests that meet the conditions (match the whitelist) within a time window. This allows for the acquisition of response data at multiple time points, thus creating multiple mock snapshots.
[0027] In this embodiment, the system continuously collects the order details interface response of order 10001 between January 1, 2026 and January 6, 2026, and obtains response data at multiple time points.
[0028] S12. Perform desensitization processing on the sensitive fields in each collected response data to obtain multiple desensitized data; For each collected response data (usually in JSON format), a recursive traversal algorithm is initiated to scan its entire tree structure, yielding multiple fields. Based on a built-in sensitive field rule base (containing multiple rules, each for example, matching idCard, mobile, password, and userPhone), sensitive field de-identification is performed. Specifically, when the name of a scanned field matches any rule in the sensitive field rule base, the de-identification algorithm corresponding to that rule (e.g., retaining the first 3 and last 4 digits of a mobile phone number, retaining the first and last digits of an ID card number) is invoked to mask or replace the field value, strictly preserving the original data type (String, Number, etc.). This is a crucial step in complying with data security regulations (such as GDPR and the Personal Information Protection Act).
[0029] In this embodiment, taking the response data obtained from a single collection as an example, the response data contains a field named userPhone, and the system anonymizes the value of this field to 138****0000.
[0030] S13. Encapsulate each de-identified data along with the corresponding collected metadata to form multiple Mock snapshots corresponding to the current interface; where the collected metadata includes: collection timestamp.
[0031] Each de-identified data point is encapsulated along with its corresponding collected metadata to form multiple mock snapshots. The collected metadata includes at least a collection timestamp. In this case, a single mock snapshot includes the collection timestamp of the corresponding response data and the de-identified data after processing. It can also be extended to include source environment identifiers, request parameter hash digests, tracking IDs, etc.
[0032] In this embodiment, the following multiple Mock snapshots are generated (examples): Mock Snapshot 1 (Collection timestamp t1=2026-01-01 10:00:00) { "orderId":"10001", "amount":99.00, "status":"PAID", "userPhone":"138****0000", "_meta":{"ts":"2026-01-01 10:00:00"} } Mock Snapshot 2 (t2=2026-01-03 09:30:00) { "orderId":"10001", "amount":109.00, "status":"PAID", "userPhone":"138****0000", "_meta":{"ts":"2026-01-03 09:30:00"} } Mock Snapshot 3 (t3=2026-01-06 15:20:00) { "orderId":"10001", "amount":109.00, "status":"REFUNDING", "userPhone":"138****0000", "_meta":{"ts":"2026-01-06 15:20:00"} } It should be noted that the _meta field is for illustrative purposes only. In actual implementation, the collected metadata can be stored separately from the data body or stored as a bypass field.
[0033] Through the above steps, the system generates multiple mock snapshots at various times on the current interface, laying the foundation for subsequent generation of time version indexes by timestamp and selection of the latest snapshot.
[0034] S2. Parse and extract the skeleton structure for each Mock snapshot, and calculate the unique contract digest hash value corresponding to each Mock snapshot by using a hash algorithm for each extracted skeleton structure; For each Mock snapshot, the skeleton structure is parsed and extracted. The skeleton structure may include a set of field names, field hierarchy, field types, object / array nesting relationships, etc. The skeleton structure emphasizes "structure rather than value".
[0035] Specifically, the process of parsing and extracting the skeleton structure of each mock snapshot essentially involves stripping away the specific field values and retaining only the data's shape information. In practice, the snapshot data (e.g., JSON format) is recursively traversed, processed according to the type of each node (object, array, primitive type): for objects, all field names are recorded, and the value of each field is recursively traversed; for arrays, the type structure of the array elements is recorded; for primitive types (string, number, boolean), their data type is recorded. After traversal, all recorded structural information is combined into a type description object, which is the skeleton structure of that snapshot.
[0036] The meanings of the various parts in the skeleton structure are as follows: the field name set refers to all the key names contained in the data, reflecting the attributes of the data returned by the interface; the field hierarchy describes the nesting depth and position of the fields in the JSON structure; the field type refers to the data type corresponding to each field (such as String, Number, Boolean, Object, Array, etc.); the object / array nesting relationship describes the containment relationship between objects and the structural definition of the elements inside the array. For example, in the three snapshots provided by the user, although the field values have changed, their field name sets (orderId, amount, status, userPhone), field types (all String or Number), field hierarchy (all top-level fields), and nesting relationship (no deep nesting) are completely consistent, therefore the skeleton structures of the three snapshots are the same.
[0037] In this embodiment, Mock snapshot 1, Mock snapshot 2, and Mock snapshot 3 have the same skeleton structure, which can be represented as follows (illustrated): orderId: String amount: Number status: String userPhone: String Metadata fields collected are not necessarily included in the skeleton structure (configurable). To more reliably identify the interface structure version, collected metadata is usually excluded from the contract skeleton.
[0038] In an optional implementation, the unique contract digest hash value corresponding to each Mock snapshot is calculated by a hash algorithm for each extracted skeleton structure, including: For each extracted skeleton structure, sort all field names in lexicographical order, and generate a normalized string corresponding to each skeleton structure based on the sorted structure; A hash algorithm is used to calculate the unique contract digest hash value corresponding to each Mock snapshot; Inject a business tag into each of the Mock snapshots.
[0039] Specifically, for each extracted skeleton structure, all field names (including those within nested objects) are recursively sorted lexicographically to eliminate differences in the original field order. Then, the sorted structure is converted into a deterministic string representation using a unified serialization format: each field is represented as "field name:data type", with multiple fields at the same level separated by commas. For nested objects, their internal fields are enclosed in curly braces and the same rules are recursively applied (e.g., userInfo:{level:String,name:String}). For arrays, element types or element structures are enclosed in square brackets (e.g., items:[{quantity:Number,skuId:String}]). In this way, skeleton structures with identical structures will generate completely identical normalized strings. For example, in Mock snapshots 1, 2, and 3 of this embodiment, after the fields are arranged lexicographically as amount, orderId, status, and userPhone, the serialized string is "amount:Number, orderId:String, status:String, userPhone:String", providing deterministic input for subsequent hash calculations.
[0040] The normalized string above is calculated using a cryptographic hash function (such as SHA-256) to generate a fixed-length hash value (e.g., a 64-bit hexadecimal string), which is the contract digest hash value. This hash value uniquely represents the structural contract of the interface corresponding to that snapshot at the time of acquisition. This hash value will change whenever the fields, types, or nesting relationships of the interface change.
[0041] For Mock snapshot 1, Mock snapshot 2, and Mock snapshot 3, since the skeleton structure is the same, the contract digest hash value is the same, and it is denoted as H_A.
[0042] To improve search accuracy, business tags can be injected into each mock snapshot. Business tags can come from any one or more of the following three categories: ① Pre-configured scene tags: Fixed tags predefined in the system or test scripts to identify the business scene to which the snapshot belongs (e.g., scene=order_detail indicates the order details scene). These tags do not change with the snapshot content and are suitable for categorization management; ② Status labels derived from field values: These labels are automatically generated based on the current values of specific fields in the snapshot response data, reflecting the business status of the snapshot (e.g., status=PAID indicates the order has been paid). These labels are dynamically updated as the data changes, facilitating retrieval by status; ③ Test configuration-specified tags: Tags temporarily specified by the test script or configuration file when running specific test cases, used to temporarily override or supplement the tag information of the snapshot (e.g., testcase=refund_test). These tags are used for flexible filtering of test dimensions.
[0043] In this embodiment, the business tag status:PAID can be injected into Mock snapshot 1 and Mock snapshot 2, the business tag status:REFUNDING can be injected into snapshot 3, and scene:order_detail can be injected uniformly.
[0044] These business tags, together with the contract summary hash, form a composite index of the snapshot.
[0045] S3. Using the contract digest hash value as the first-level index and the collection timestamp as the second-level index, store each Mock snapshot containing two levels of indexes into the snapshot database; In this invention, a storage model is designed within the snapshot database. The first-level index (primary dimension) is the contract digest hash value, which aggregates all mock snapshots with identical skeleton structures into the same logical group, representing a "semantic version" of the current interface. The second-level index (secondary dimension) is the collection timestamp, which arranges all mock snapshots in chronological order within the logical group of the same semantic version, clearly showing how the response data evolves over time during a period when the current interface structure is stable. This storage method can be implemented using various forms, such as relational database table structures, key-value databases (first-level key is the contract digest hash value, second-level key is the collection timestamp), and document database indexes.
[0046] Furthermore, business tags are not independent index levels, but rather stored as supplementary metadata for snapshots (e.g., saved along with snapshot data, or as a bypass field). In the two-level index-based storage model, queries first locate the corresponding logical group using the first-level index (contract digest hash value), then traverse or locate snapshots within that logical group using the second-level index (collection timestamp), and filter based on business tags during the traversal. To optimize tag query performance, auxiliary indexes (such as inverted indexes) can be created for business tags to map tag values to a snapshot ID list, but the core storage model still maintains a two-level structure based on contract digests and timestamps, ensuring efficient structure version management and time-series retrieval.
[0047] An example storage structure is shown below (for illustration only): First-level index: H_A o Secondary index: t1 -> S1 o Secondary index: t2 -> S2 o Secondary index: t3 -> S3 With the help of a two-level index, the system can perform the following common queries: Given a structure signature H_A, retrieve the latest snapshot in reverse timestamp order (e.g., Mock snapshot 3). Given a structure signature H_A and a business label (e.g., status=PAID), retrieve the latest snapshot (e.g., Mock snapshot 2) from the snapshots that meet the label conditions by timestamp. Given a set of structure signatures, select from multiple structure versions according to a strategy.
[0048] This index structure provides the foundation for subsequent IR-based retrieval of target snapshots.
[0049] In an optional implementation, the method further includes the following after S3: Whenever a new Mock snapshot is generated (e.g., when Mock snapshot 2 or Mock snapshot 3 is generated), a hierarchical difference comparison is performed between the new Mock snapshot and the old Mock snapshot of the previous time version of the current interface (e.g., Mock snapshot 1). The hierarchical difference comparison includes at least a structural layer difference comparison and a key field value layer difference comparison. Output a difference report based on the comparison results and mark the nature of the differences.
[0050] (1) Comparison of structural layer differences Structural layer difference comparison: If the contract digest hash value of the new snapshot differs from that of the old snapshot, a detailed analysis of the skeleton structure differences (such as structural changes like field additions, field deletions, and field type changes) is performed. When the interface evolves, mock snapshots from different times may belong to different contract digest hash values. Structural layer difference comparison can be used for: Confirm that the structure of the new snapshot is consistent with the existing structure; If there is a discrepancy, locate the difference field and difference type, and mark the nature of the difference (such as compatibility change / destructive change).
[0051] (2) Comparison of differences in key field values Value layer difference comparison: If the contract digest hash values are the same, it indicates that the skeleton structures of the two are consistent. In this case, value layer difference comparison is performed on key business fields to analyze the specific changes in field values. Key business fields can be determined by pre-configuring a list of key fields (such as amount, status, and other core business fields) in the system. For example, for the order interface, the default values are amount and status as key business fields. When the amount in Mock snapshot 1 is 99.00 and the amount in Mock snapshot 2 becomes 109.00, the value layer difference comparison will identify the change and record it.
[0052] (3) Difference reporting and difference nature marking The contents and meanings of the difference report are as follows: Difference path: refers to the specific location of the changed field in the JSON structure. For example, / payment_method represents the payment_method field under the root node. If it is a nested structure, the levels are separated by forward slashes (such as / user / address / city). Difference type: refers to the nature of the change, including field addition (not present in the original structure), field missing (present in the original structure but absent in the current structure), field type change (e.g., String becomes Number), etc. Difference nature: Used to assess the degree of impact of changes on the interface contract. Compatibility changes (such as adding optional fields) usually do not affect older version callers, while destructive changes (such as deleting fields or changing types) may cause older version calls to fail. Scope of impact: This refers to the range of test cases or business modules that may be affected by the difference.
[0053] The difference attribute flag can be used to select a test strategy for the current interface. If the difference attribute flag is "Compatibility New Field", the system can automatically fill in the default value for the snapshot and continue to use the snapshot to execute tests without manual intervention. If it is flagged as "Destructive Type Change", the system can trigger a degradation strategy, automatically select the old structure snapshot before the change for playback, and prompt the test cases to be upgraded to adapt to the new interface. In this way, the test process can adaptively adjust according to the compatibility of interface changes, reducing manual intervention and test interruptions.
[0054] In this embodiment, the system will demonstrate the role of this step when subsequent structural changes occur.
[0055] S4. Before executing the test using the test case, obtain the intermediate representation IR of the test case, and retrieve the target snapshot in the snapshot database corresponding to the current interface based on the intermediate representation IR and the two-level index; wherein, when a change is detected in the current interface, the target snapshot is retrieved again according to the changed current interface; In one alternative implementation, "the current interface has changed" can be triggered in various ways, such as changes to the current interface definition file (meaning that the specification file describing the current interface contract (such as OpenAPI, Swagger file) has been updated, indicating that the current interface definition itself has been modified); structural layer difference report prompts, etc.; the present invention does not limit the triggering method.
[0056] For example: The current interface has added a "payment method (payment_method)" field.
[0057] Intermediate representations are abstract descriptions of test cases, generated by parsing test case code or its configuration files. Based on IR (In-Reference Function) for target snapshot retrieval, the retrieval process becomes "use case requirement-oriented": instead of simply selecting the latest or random snapshot, it prioritizes snapshots that structurally fully cover the fields required by the test case, and whose field types match the assertion expectations. This approach significantly improves replay success rate (avoiding test failures due to missing fields or type mismatches) while enhancing interpretability (retrieval results can be directly linked to the specific requirements of the test case).
[0058] Before a testing framework (such as JUnit or pytest) executes a test case, the system parses the test case code or its configuration file to generate an intermediate representation (IR). From the IR, the explicit dependencies of the test case are extracted: including the path to the interface to be called, the request parameters used, and the response fields referenced in the assertions and their expected types.
[0059] In an optional implementation, retrieving the target snapshot from the snapshot database corresponding to the current interface based on the intermediate representation IR and utilizing the two-level index includes: Based on the intermediate representation IR, and using the two-level index, multiple candidate snapshots are retrieved in the snapshot database corresponding to the current interface; Calculate the matching score between each candidate snapshot and the test case, and select the candidate snapshot with the highest score as the target snapshot; Perform a contract consistency check on the target snapshot; if the check fails, trigger a degradation strategy to select an alternative snapshot as the target snapshot or output an exception.
[0060] Based on the intermediate representation IR and utilizing the two-level index, multiple candidate snapshots are retrieved in the snapshot database corresponding to the current interface. The specific steps are as follows: First, the dependency information of the test cases is parsed from the IR, including the identifier of the invoked interface (determined by the URL path pattern and HTTP method), request parameters, and response fields referenced in the assertions and their expected types. Based on this information, the expected skeleton structure characteristics (i.e., the set of fields and type constraints) can be derived for subsequent structure matching.
[0061] Based on this, the initial range of candidate snapshots is determined using one of the following two methods: Method 1 (Skeleton Structure Pre-screening): Calculate the corresponding contract digest hash value (which may be one or more) based on the expected skeleton structure features. Use the first-level index (contract digest hash value) to locate the logical group with the same skeleton structure in the snapshot database. Then, within this group, use the second-level index (collection timestamp) to obtain a set of snapshots (e.g., take the N most recent ones in reverse chronological order) as candidate snapshots.
[0062] Method 2 (Time Priority): If pre-screening is not based on the skeleton structure, the most recent few Mock snapshots (e.g., the first K in reverse timestamp order) are directly obtained from the snapshot database of the current interface using the second-level index as candidate snapshots.
[0063] After obtaining the initial candidate snapshots, the candidate range can be further narrowed down by combining business tags: by comparing the business tags of each candidate snapshot with the scene tags implicit in the IR (such as tags extracted from test case names, configurations, or comments, such as scene=order_detail) or dynamically inferred tags (such as states inferred from request parameters), snapshots with matching tags are filtered out. If no tag information is provided in the IR, all initial candidate snapshots are retained.
[0064] Ultimately, a set of candidate snapshots is obtained, which are used for subsequent matching score calculations (such as combining contract compatibility, tag matching degree, and time freshness) and selection of target snapshots. This process fully utilizes the positioning capabilities of the two-level index and improves the accuracy of retrieval through tag filtering.
[0065] Calculate the matching score between each candidate snapshot and the test case, and select the candidate snapshot with the highest score as the target snapshot; The matching score is obtained by weighting the contract compatibility score, the tag matching score, and the time freshness score. Contract Compatibility Score (Highest Weight): Based on the set of fields and their types required for the test cases parsed from the intermediate representation IR, this score checks whether the skeleton structure of the candidate snapshot fully contains these fields and their types match. If fully satisfied, the score is 1 (or full marks); otherwise, the score is 0 (or a score calculated based on the field coverage ratio). This score has the highest weight and is used to ensure that the snapshot structure meets the basic requirements of the test cases.
[0066] Tag matching score (medium weight): The business tags of the candidate snapshots are compared with the expected tags of the test cases (scene tags that can be extracted from the test case name, configuration file or code comments, such as scene=order_detail). The score is obtained by calculating the tag overlap (such as the proportion of the number of common tags to the total number of expected tags). The weight is medium.
[0067] Time freshness score (lowest weight): The score is calculated by normalizing the time difference between the collection timestamp of the candidate snapshot and the current time (e.g., using negative exponential decay or linear normalization). The closer the time is to the current moment, the higher the score and the lowest weight.
[0068] The system calculates a comprehensive score for each candidate snapshot and selects the snapshot with the highest score as the target snapshot.
[0069] After selecting the target snapshot, a contract consistency check is performed to ensure that the target snapshot meets the dependency requirements of the test cases. The check is based on the test case information parsed from the intermediate representation (IR), mainly including three aspects: field existence, field type, and field value range. For example: Field Existence Validation: Verifies whether the target snapshot contains the fields that the test case depends on. If the IR indicates that the test case expects a certain field to exist (e.g., for assertions or post-processing), then the target snapshot must contain that field; if the test case allows for a missing field and can be filled in with a default value, then it is considered to pass. For example, if the test case assertion references the `payment_method` field, the validation passes if the target snapshot contains this field; if the target snapshot lacks this field, but the system configuration allows for automatic filling in of a default value for this field (e.g., an empty string or the default value "UNKNOWN"), then it is still considered compatible; otherwise, the validation fails.
[0070] Field type validation: Verifies whether the data type of a specific field in the target snapshot is consistent with the type declared in the IR. For example, if the IR requires the amount field to be a numeric type for calculations, then the value of amount in the target snapshot must be a number (such as 109.00); if it is a string type (such as "109.00"), then the type mismatch will fail and the validation will fail.
[0071] Field value range validation: Verifies whether the value of a specific field in the target snapshot is within the enumeration range allowed by IR. For example, if IR requires that the value of the status field can only be "PAID" or "REFUNDING", then the value of status in the target snapshot must belong to this set; if the value is "CANCELLED", the validation will fail.
[0072] When verification fails, a degradation strategy is triggered, and the options include: Select the candidate snapshot with the second highest score as the target snapshot; Switch to the old structure version (i.e., a snapshot of the same interface that had a compatible structure before this interface change, usually the last version before the change). The output shows an error and prompts that the snapshot or test case needs to be updated.
[0073] For example, if Mock snapshots 1, 2, and 3 do not have a `payment_method` field, but other core fields match, and this new field allows being empty or having a default value, the system determines it to be "forward compatible." The latest Mock snapshot, 3, is then provided to the test code.
[0074] When a change to the current interface is detected, the above retrieval process will be re-executed based on the changed current interface definition to ensure that the target snapshot reflects the latest current interface contract.
[0075] S5. During test execution, the target snapshot is used as the response, and the real response is acquired in parallel for difference comparison and output of difference view; stable difference identification is performed on the difference between the real response and the target snapshot, and a patch is generated for the identified stable difference. The patch is applied to the target snapshot to generate a new time version of the Mock snapshot.
[0076] In an optional implementation, when a test case initiates a network request to the current interface, the replay control module (e.g., via an HTTP client interceptor or a WireMock server) intercepts the request. This module matches the request characteristics (URL, parameters) to a pre-selected target snapshot, then dynamically constructs an HTTP response, using the anonymized response data from the target snapshot as the response body, and sets the appropriate HTTP status code and response headers, returning it directly to the test case. This process completely blocks network calls to the real service, ensuring the independence and replayability of the test.
[0077] In one implementation, without affecting replay (i.e., by initiating requests asynchronously or in parallel threads without blocking the return of mock responses), a request identical to the test case is sent to the real service environment corresponding to the current interface (such as a real business service in a production, pre-production, or test environment) to obtain the real response (i.e., the real response data). The obtained real response is compared with the target snapshot, and a difference view is output. The difference view may include: missing fields, added fields, different field types, changes in key field values, etc., used to locate problems and guide snapshot evolution.
[0078] For example: The test code actually reads `payment_method`, but it doesn't exist in the S3 snapshot. Solution: The system generates a difference view, explicitly indicating that "the data is lagging behind the code".
[0079] Stable difference identification is performed on the difference between the actual response and the target snapshot. This includes: Noise filtering is performed on the difference between the actual response and the target snapshot to exclude pre-defined non-business field differences; the remaining differences after filtering are statistically analyzed, and when the same difference appears in N consecutive test executions of the same interface, it is determined to be a stable difference.
[0080] Noise filtering here refers to automatically excluding non-business-related differences during the comparison process. The noise filter identifies and ignores these fields in the following ways: Field name matching: Based on a preset keyword library or regular expressions, identify and exclude fields that usually contain non-business data (such as timestamp, UUID, token, etc.). Field value pattern matching: For fields that cannot be identified by field name alone, identification is performed by value format features (such as UUID format, timestamp format); Field path specification: Precisely specify the location of the field to be ignored using JSON path expressions.
[0081] For example, the actual response and the target snapshot may differ in the timestamp and requestId fields, but these two fields are identified and excluded by the noise filter and are not included in the subsequent stable difference statistics.
[0082] For any remaining suspected business field differences after filtering (such as status changing from 3 to 4), the system will record their frequency. If the same difference in the same field occurs in N consecutive tests (e.g., N=3) targeting the same interface, it is determined to be a stable difference, indicating that this is likely a formal change in interface behavior or business logic, rather than an occasional phenomenon.
[0083] JSON patches are generated for identified stable differences. A JSON patch is a standard format used to describe changes to a JSON document, consisting of a series of operations, each including an operation type (op), a target path (path), and an optional value (value). In the specific implementation, the corresponding operation is generated based on the type of each stable difference. If a new field is added in the actual response but the target snapshot is missing, then an add operation is generated; If the type or value undergoes a stable change, a replace operation is generated; If the field is deleted stably, a remove operation is generated.
[0084] The ordered set of these operations constitutes a JSON patch that fully describes the changes from the target snapshot to the actual response.
[0085] After generating the patch, the system applies it to the target snapshot, creating a new mock snapshot. The application process involves modifying the target snapshot's data structure sequentially according to the operation order in the patch to obtain the updated response data. The new mock snapshot's acquisition timestamp is set to the current system time (i.e., the moment the patch was generated and applied) to identify this snapshot version.
[0086] Subsequently, the new mock snapshot is stored in the snapshot database: First, its contract digest hash value is calculated based on the skeleton structure of the new mock snapshot, serving as the first-level index; then, the collection timestamp of the new mock snapshot is used as the second-level index, and the new mock snapshot is stored in the corresponding hash group. If the skeleton structure of the new mock snapshot differs from the original target snapshot (i.e., the patch caused a structural change), its contract digest hash value will also change. In this case, the new mock snapshot will be stored in a new hash group, while the original target snapshot remains in the original group. This mechanism ensures that the original target snapshot is never overwritten, thus fully preserving historical versions and supporting subsequent backtracking and auditing.
[0087] For example: Based on Mock snapshot 3, the default value payment_method: "UNKNOWN" is automatically added to generate Mock snapshot 4.
[0088] While generating a new snapshot, an audit log is recorded, including: the test case ID that triggered the update, the change time, the application's JSON patch, the business fields affected by the change and their previous and current values, and an inference about the reason for the change. This immutable chain of evidence is stored in the snapshot database along with the new snapshot.
[0089] This completes a full governance loop of "collection -> indexing -> matching -> replay -> comparison -> update". The Mock data pool can continuously, automatically, and accurately update in sync with the evolution of the real interface, always maintaining its timeliness and availability.
[0090] Figure 2 This is a schematic diagram of the structure of a Mock data lifecycle governance and replayable system provided in an embodiment of the present invention; as shown below. Figure 2 As shown, the system includes: The snapshot generation unit 201 is used to listen to requests to the current interface that hit the preset whitelist, collect the response data of each request, perform de-identification processing on each response data, and generate multiple Mock snapshots corresponding to the current interface. The hash calculation unit 202 is used to parse each of the Mock snapshots and extract the skeleton structure, and calculate the unique contract digest hash value corresponding to each Mock snapshot by using a hash algorithm for each extracted skeleton structure. Snapshot storage unit 203 is used to store each Mock snapshot containing two levels of indexes to the snapshot database, with the contract digest hash value as the first level index and the collection timestamp as the second level index. The target snapshot retrieval unit 204 is used to obtain the intermediate representation IR of the test case before executing the test using the test case, and retrieve the target snapshot in the snapshot database corresponding to the current interface based on the intermediate representation IR and the two-level index; wherein, when a change is detected in the current interface, the target snapshot is retrieved again according to the changed current interface; The evolution and audit unit 205 is used to take the target snapshot as a response during test execution, and obtain the real response in parallel for difference comparison and output a difference view; to identify stable differences between the real response and the target snapshot, generate patches for the identified stable differences, and apply the patches to the target snapshot to generate a new time version of the Mock snapshot.
[0091] In an optional implementation, the snapshot generation unit 201 includes: The matching subunit 2011 is used to listen to network requests for the current interface. It combines the requested URL path and HTTP method as a combination and performs regular expression matching with a preset whitelist. If the match is successful, it continuously collects response data for the current interface at preset intervals. If the match is unsuccessful, it allows the request to proceed without data collection. The desensitization subunit 2012 is used to perform desensitization processing on sensitive fields in each collected response data to obtain multiple desensitized data. The encapsulation subunit 2013 is used to encapsulate each de-identified data along with the corresponding collected metadata to form multiple Mock snapshots corresponding to the current interface; among which, the collected metadata includes: collection timestamp.
[0092] In an optional implementation, the step of calculating the unique contract digest hash value corresponding to each Mock snapshot using a hash algorithm for each extracted skeleton structure includes: For each extracted skeleton structure, sort all field names in lexicographical order, and generate a normalized string corresponding to each skeleton structure based on the sorted structure; A hash algorithm is used to calculate the unique contract digest hash value corresponding to each Mock snapshot for each normalized string. Inject a business tag into each of the Mock snapshots.
[0093] In an optional implementation, the system further includes: a difference comparison unit, used for: Whenever a new Mock snapshot is generated, a hierarchical difference comparison is performed between the new Mock snapshot and the old Mock snapshot of the previous time version of the current interface. The hierarchical difference comparison includes at least a structural layer difference comparison and a key field value layer difference comparison. Output a difference report based on the comparison results and mark the nature of the differences.
[0094] The system of the present invention corresponds to the method described above, and the specific implementation of the system will not be repeated here.
[0095] In summary, this invention continuously collects and de-identifies data from the current whitelisted interfaces to generate multiple mock snapshots; extracts the skeleton structure from the snapshots and normalizes the hash to obtain the contract digest hash value; constructs a two-level index storage using the contract digest hash value and the collection timestamp; retrieves the target snapshot based on IR and replays it before interface changes and test execution to achieve replayability; and optionally outputs a difference view for real comparison. Through stable difference identification and JSON patch generation, the snapshots are minimized and their evolution is minimized while recording audit logs, thus forming a continuous governance closed loop for mock snapshots. This mechanism effectively solves the problems of easy aging of traditional static mocks, non-replayability in real environments, and lack of version governance and auditing, improving test stability, reproducibility, and governance efficiency.
[0096] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for Mock data lifecycle governance and replayability, characterized in that, include: S1. Listen for requests to the current interface that hit the preset whitelist, collect the response data of each request, perform de-identification processing on each response data, and generate multiple Mock snapshots corresponding to the current interface. S2. Parse and extract the skeleton structure for each Mock snapshot, and calculate the unique contract digest hash value corresponding to each Mock snapshot by using a hash algorithm for each extracted skeleton structure; S3. Using the contract digest hash value as the first-level index and the collection timestamp as the second-level index, store each Mock snapshot containing two levels of indexes into the snapshot database; S4. Before executing the test using the test case, obtain the intermediate representation IR of the test case, and retrieve the target snapshot in the snapshot database corresponding to the current interface based on the intermediate representation IR and the two-level index; wherein, when a change is detected in the current interface, the target snapshot is retrieved again according to the changed current interface; S5. During test execution, the target snapshot is used as the response, and the real response is acquired in parallel for difference comparison and output of difference view; stable difference identification is performed on the difference between the real response and the target snapshot, and a patch is generated for the identified stable difference. The patch is applied to the target snapshot to generate a new time version of the Mock snapshot.
2. The method according to claim 1, characterized in that, S1 includes: Listen for network requests to the current interface, combine the requested URL path and HTTP method into a single string, and perform regular expression matching against a preset whitelist. If a match is found, continuously collect response data from the current interface at preset intervals. If a match is not found, allow the request directly without collecting any data. Sensitive fields in each collected response data are de-identified, resulting in multiple de-identified data sets. Each de-identified data point is encapsulated together with its corresponding collected metadata to form multiple Mock snapshots for the current interface; the collected metadata includes: collection timestamp.
3. The method according to claim 1, characterized in that, The step of calculating the unique contract digest hash value corresponding to each Mock snapshot by using a hash algorithm for each extracted skeleton structure includes: For each extracted skeleton structure, sort all field names in lexicographical order, and generate a normalized string corresponding to each skeleton structure based on the sorted structure; A hash algorithm is used to calculate the unique contract digest hash value corresponding to each Mock snapshot; Inject a business tag into each of the Mock snapshots.
4. The method according to claim 1, characterized in that, Following S3, the following is also included: Whenever a new Mock snapshot is generated, a hierarchical difference comparison is performed between the new Mock snapshot and the old Mock snapshot of the previous time version of the current interface. The hierarchical difference comparison includes at least a structural layer difference comparison and a key field value layer difference comparison. Output a difference report based on the comparison results and mark the nature of the differences.
5. The method according to claim 1, characterized in that, The step of retrieving the target snapshot from the snapshot database corresponding to the current interface based on the intermediate representation IR and utilizing the two-level index includes: Based on the intermediate representation IR, and using the two-level index, multiple candidate snapshots are retrieved in the snapshot database corresponding to the current interface; Calculate the matching score between each candidate snapshot and the test case, and select the candidate snapshot with the highest score as the target snapshot; Perform a contract consistency check on the target snapshot; if the check fails, trigger a degradation strategy to select an alternative snapshot as the target snapshot or output an exception.
6. The method according to claim 5, characterized in that: The matching score is obtained by weighting the contract compatibility score, tag matching score, and time freshness score.
7. The method according to claim 1, characterized in that, The stable difference identification of the difference between the true response and the target snapshot includes: Noise filtering is performed on the difference between the actual response and the target snapshot to exclude pre-defined non-business field differences; the remaining differences after filtering are statistically analyzed, and when the same difference appears in N consecutive test executions of the same interface, it is determined to be a stable difference.
8. A Mock data lifecycle governance and replayability system, characterized in that, include: The snapshot generation unit is used to listen to requests to the current interface that hit the preset whitelist, collect the response data of each request, perform de-identification processing on each response data, and generate multiple Mock snapshots corresponding to the current interface. The hash calculation unit is used to parse each Mock snapshot and extract the skeleton structure, and calculate the unique contract digest hash value corresponding to each Mock snapshot by using a hash algorithm for each extracted skeleton structure. A snapshot storage unit is used to store each Mock snapshot containing two levels of indexes into a snapshot database, with the contract digest hash value as the first-level index and the collection timestamp as the second-level index. The target snapshot retrieval unit is used to obtain the intermediate representation IR of the test case before executing the test case, and retrieve the target snapshot in the snapshot database corresponding to the current interface based on the intermediate representation IR and the two-level index; wherein, when a change is detected in the current interface, the target snapshot is retrieved again according to the changed current interface; The evolution and auditing unit is used to take the target snapshot as a response during test execution, and simultaneously acquire the real response for difference comparison and output a difference view; to identify stable differences between the real response and the target snapshot, generate patches for the identified stable differences, and apply the patches to the target snapshot to generate a new time version of the Mock snapshot.
9. The system according to claim 8, characterized in that, The snapshot generation unit includes: The matching subunit is used to listen to network requests for the current interface. It combines the requested URL path and HTTP method as a combination and performs regular expression matching against a preset whitelist. If the match is successful, it continuously collects response data for the current interface at preset intervals. If the match is unsuccessful, it allows the request to proceed without data collection. The desensitization subunit is used to perform desensitization processing on sensitive fields in each collected response data to obtain multiple desensitized data. The encapsulation subunit is used to encapsulate each de-identified data along with the corresponding collected metadata to form multiple Mock snapshots corresponding to the current interface; among which, the collected metadata includes: collection timestamp.
10. The system according to claim 8, characterized in that, The step of calculating the unique contract digest hash value corresponding to each Mock snapshot by using a hash algorithm for each extracted skeleton structure includes: For each extracted skeleton structure, sort all field names in lexicographical order, and generate a normalized string corresponding to each skeleton structure based on the sorted structure; A hash algorithm is used to calculate the unique contract digest hash value corresponding to each Mock snapshot for each normalized string. Inject a business tag into each of the Mock snapshots.