A property insurance loss reporting data semantic alignment method based on multi-source query feedback

CN122817364APending Publication Date: 2026-09-25HEFEI GUOKE DIGITAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611162163.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-03
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]本发明的实施例提供一种基于多源查询反馈的财产保险报损数据语义对齐方法,用于解决由于标准项目语义体系与各理赔数据源的记录口径及覆盖范围不一致,导致仅按语义匹配分值确定的标准项目发生多源检索错配的问题

Benefits of technology

1、本发明基于语义匹配结果进行语义检索,得到候选标准项目,不直接把语义匹配分值最高的候选标准项目作为最终结果,而是将各候选标准项目分别转换为多源试探查询,并以实际返回记录对语义候选进行验证,使标准项目的确定同时受到文本语义和数据源实际记录的约束,从而减少标准项目与后续可检索记录不对应造成的检索错配。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817364A_ABST
    Figure CN122817364A_ABST
Patent Text Reader

Abstract

The application provides a property insurance loss reporting data semantic alignment method based on multi-source query feedback, which is applied to the technical fields of heterogeneous database query and semantic matching, and the method performs semantic matching on property insurance loss reporting item data and standard item data, performs semantic retrieval on the standard item data based on the semantic matching result, and generates candidate standard items; each candidate standard item is converted into a heuristic query of different claim data sources, data source unavailability, category non-coverage and effective no hit are distinguished, and a support state, an opposing state or a neutral state is formed according to the number of hit records, record content consistency and item granularity consistency, multi-source feedback support data is generated by summarizing, the query is corrected in a targeted manner when the feedback correction condition is met, and the target standard item is determined in combination with the semantic matching score and the updated feedback support.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of heterogeneous database query and semantic matching technology, and in particular to a semantic alignment method for property insurance loss reporting data based on multi-source query feedback. Background Technology

[0002] In the property insurance claims process, on-site photos, customer loss reports, and inspection records can be manually entered, and processed through image recognition or text recognition to form a digital loss report containing the item name, quantity, amount, and attributes of the damaged object. To enable the loss report items to be linked to price records and historical compensation records, it is usually necessary to map the non-standard names used by customers to unified standard item names or standard classification codes.

[0003] Existing semantic alignment methods typically extract the name and context of the reported claim item, perform literal matching, semantic vector matching, and hierarchical matching with standard names in standard item libraries such as UNSPSC, sort according to matching scores, select the standard item with the highest score or that reaches the threshold, and then use the standard item to query the self-built claims item database, historical claims case database, or market inquiry data source.

[0004] However, each claims data source is independently formed, and their field structures, item naming, record granularity, and category coverage are inconsistent. For example, the higher-level comprehensive items in the standard item library are semantically close to the description of the loss, while historical claims records and market records may only be stored according to specific materials or repair sub-items. In this case, the standard item with the highest semantic score may not correspond to the actual record in each data source, which can easily lead to valid but unmatched results or the return of irrelevant records. Therefore, because the semantic system of standard items is inconsistent with the record scope and coverage of each claims data source, multi-source retrieval mismatches occur due to standard items determined solely by semantic matching scores. Summary of the Invention

[0005] This invention provides a semantic alignment method for property insurance claim data based on multi-source query feedback. This method addresses the problem of multi-source retrieval mismatches in standard items determined solely by semantic matching scores due to inconsistencies in the semantic system of standard items and the recording scope and coverage of various claim data sources. To achieve the above objective, this invention employs the following technical solution: A semantic alignment method for property insurance claim data based on multi-source query feedback includes: acquiring property insurance claim item data, standard item data, and claims data source configuration data, wherein the claims data source configuration data includes query field mapping relationships, data coverage, and data source availability status; performing semantic matching between property insurance claim item data and standard item data, and performing semantic retrieval on the standard item data based on the semantic matching results to obtain candidate standard item data containing semantic matching scores; converting each candidate standard item data into trial query data based on the query field mapping relationships, performing trial queries on available claims data sources whose data coverage includes candidate standard item categories, and performing trial queries on unavailable or... Neutral feedback is generated from claim data sources that do not cover candidate standard item categories, resulting in multi-source query feedback data. Based on query execution status, number of hit records, record content consistency, and item granularity consistency, the feedback from each claim data source to candidate standard items is divided into supportive, opposing, or neutral states, and multi-source feedback support data is generated. When the multi-source feedback support data meets preset feedback correction conditions, query correction data is generated based on the multi-source query feedback data, and a trial query is executed again to update the multi-source feedback support data. Target standard items are determined based on semantic matching scores and multi-source feedback support data, generating semantic alignment results between property insurance loss reporting item data and target standard items.

[0006] As can be seen from the above technical solution, the present invention has the following beneficial effects: 1. This invention performs semantic retrieval based on semantic matching results to obtain candidate standard items. Instead of directly taking the candidate standard item with the highest semantic matching score as the final result, each candidate standard item is converted into a multi-source trial query, and the semantic candidates are verified by the actual returned records. This ensures that the determination of standard items is constrained by both text semantics and the actual records of the data source, thereby reducing retrieval mismatches caused by the mismatch between standard items and subsequent searchable records.

[0007] 2. This invention identifies data source unavailability, query anomalies, category miscovery, and incomplete hit record fields as neutral feedback, and identifies successful queries but no hits or inconsistent complete records as negative feedback. This avoids incorrectly reducing the support level of candidate standard items due to the data source lacking effective evaluation conditions. Targeted re-queries are only performed on candidate items and data sources that meet the feedback correction conditions, which can reduce repeated full queries and irrelevant record transmissions.

[0008] 3. This invention identifies one-to-one, one-to-many, or many-to-one alignment relationships based on the consistency of record content and project granularity in the query feedback, and retains source identifiers and query evidence, so that the semantic alignment results can stably connect to the self-built claims project database, historical claims case database, and market inquiry data source, providing traceable standardized data for subsequent loss assessment or valuation. Attached Figure Description

[0009] The invention will now be further described with reference to the accompanying drawings.

[0010] Figure 1 The overall flowchart of a semantic alignment method for property insurance loss reporting data based on multi-source query feedback provided by the present invention; Figure 2 This invention provides a flowchart for the generation of multi-source query feedback and feedback support data. Figure 3 The flowchart for query correction and target standard item determination provided by this invention. Detailed Implementation

[0011] The terms "first," "second," and "third," etc., used in this specification, claims, and drawings are used to distinguish different objects, not to limit a specific order.

[0012] In embodiments of the present invention, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" or "for example" in embodiments of the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Rather, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0013] In this embodiment, property insurance claim item data refers to structured item records formed by customer claim lists, inspection records, or multimodal recognition results; standard item data refers to standard names, standard codes, and hierarchical relationships stored according to a unified item classification system; multi-source query feedback data is not manually evaluated, but rather the machine execution response of exploratory queries in different claims data sources and the neutral feedback generated for data sources that have not executed queries; valid query feedback data is feedback that the data source is available, the category is covered, and the query was successfully executed; item granularity consistency is used to indicate whether the item hierarchy of the returned records is compatible with the claim item structure; multi-source feedback support data is used to record the support status, opposition status, or neutral status of different claims data sources for the same candidate standard item and their summary level, including cross-source support level, single-source support level, no support level, source conflict level, and insufficient feedback level.

[0014] Research has revealed that while existing semantic matching can identify standard items with similar names, semantic similarity only indicates a similar textual meaning and does not prove that the standard item can be matched with the corresponding record in the subsequent claims data source. In particular, when the same comprehensive item is broken down into material items, repair items, or historical damage assessment items by different data sources, directly using the highest semantic matching result will propagate semantic selection errors to the database query stage.

[0015] To address the aforementioned issues, this invention provides a semantic alignment method for property insurance loss reporting data based on multi-source query feedback: First, multiple semantic candidates are retained, and then multi-source trial queries with limited fields and limited record counts are executed for each candidate standard item; subsequently, data source unavailability, category non-coverage, query anomalies, and valid misses are distinguished, and a source support status and multi-source feedback support level are formed based on the actual returned records. Targeted correction queries are only executed for candidates with misses, low consistency, or source conflicts, ultimately determining the target standard item corresponding to the actual record.

[0016] Example: like Figures 1 to 3 As shown, this embodiment relates to a semantic alignment method for property insurance loss reporting data based on multi-source query feedback, which is particularly suitable for scenarios where the property insurance loss reporting list needs to be associated with a self-built claims project database, a historical claims case database, and a market inquiry data source.

[0017] Specific implementation steps: S1. Obtain data on reported damage items, standard items, and claims data source configuration data; In this embodiment, step S1 includes the following process: The system retrieves property insurance claim data from the claims processing system. This data can be entered by surveyors or generated through on-site photo recognition, text recognition, and extraction of list fields. Each claim record contains at least a claim identifier and a claim name. To support subsequent disambiguation, the system also retrieves the damaged object attributes, handling action attributes, claim scenario attributes, project structure relationships, and the original claim amount. The damaged object attributes indicate the category, material, specifications, or location of the damaged object; the handling action attributes indicate actions such as repair, replacement, dismantling, or restoration; and the project structure relationships indicate whether the project is an independent project, a comprehensive project, or a divisible sub-project.

[0018] Standard item data is read from the standard item library. The standard item library can adopt the UNSPSC classification system or a claims standard item system developed by insurance institutions based on UNSPSC. Each standard item data entry includes a standard item identifier, standard item name, standard synonym expression, standard hierarchy, and standard item category. The standard hierarchy record at least the parent item identifier and child item identifier, enabling the processor to read parent items, child items, and adjacent items at the same level along the standard classification tree. Standard synonym expressions are derived from aliases in standard files, verified historical loss statements, and verified query aliases from various data sources; unverified automatic alignment results are not directly written to the standard synonym expressions.

[0019] Read the claims data source configuration data. Claims data sources must include at least two of the following: a self-built claims project database, a historical claims case database, and a market inquiry data source. Configure the data source identifier, access interface, query fields, return fields, query field mapping relationships, data coverage, data source availability status, and maximum number of query records for each claims data source. The query field mapping relationships specify which fields in the data source should contain the standard project identifier, standard project name, damaged object attribute, handling action attribute, regional attribute, and time attribute. The data coverage is determined based on the data source field directory and historical valid records, and must at least include the searchable standard project category, regional range, and time range.

[0020] The availability status of a data source is obtained through a lightweight status query before semantic alignment is performed. For database-type data sources, database connectivity, table structure versions, and access permissions can be confirmed through connection tests and preset read-only status statements; for API-type data sources, authentication and response status can be confirmed through API health checks. Status queries do not read complete business records. When a connection fails, authentication fails, or table structure version is incompatible, the data source availability status is marked as unavailable, and the reason for the status is recorded to avoid misjudging the data source as unavailable due to a lack of records in candidate standard items.

[0021] S2. Perform semantic matching and generate candidate standard item data; In this embodiment, step S2 includes the following process: First, the names of reported damage items are standardized by unifying full-width and half-width characters, removing fields such as quantity, currency, and serial number that do not contribute to the meaning of the item, and replacing internal abbreviations with confirmed standard synonyms. Then, the standardized damage item names are associated with the attributes of the damaged object, the handling action, and the damage reporting scenario to form semantic feature data for the damage items. During the association process, the source identifier of each field is retained, and handling actions such as "replacement" and "repair" are not merged with the damaged object name into a single, untraceable text.

[0022] The semantic feature data of the reported damage items and the data of each standard item are matched. The literal matching degree is obtained based on the common terms and their order; the semantic vector matching degree is calculated by converting the semantic feature data of the reported damage items and the names of the standard items into vectors using an existing text vector model; the scene matching degree is obtained based on the consistency between the attributes of the damaged object, the attributes of the handling action, and the attributes of the reported damage scene and the applicable description of the standard item; the standard level matching degree is obtained based on whether the structural relationship of the reported damage items matches the level of the standard items. All of the above are normalized to closed intervals. Then, generate the first according to the following formula. Semantic matching scores for each candidate standard item: ; In the formula, For the first Semantic matching scores of each candidate standard item; , , and The four levels of matching are, in order: literal matching degree, semantic vector matching degree, scene matching degree, and standard level matching degree, all of which take values ​​within a closed interval. Dimensionless numbers within; , , and The following are the weights of the corresponding matching items, all of which are non-negative dimensionless numbers and satisfy the following conditions: ; This refers to the candidate standard item number.

[0023] The weights and candidate thresholds for each matching item are determined by historically confirmed semantic mapping records. Specifically, confirmed correct mappings are used as positive samples, and other standard items not confirmed in the same matching are used as negative samples. The aforementioned matching items are calculated for each sample. Candidate thresholds are selected sequentially from the actual scores of the samples, and the number of correct candidates retained and incorrect candidates excluded is statistically analyzed in the reserved verification samples. The threshold that meets the preset retention rate condition and has a small number of incorrectly retained items is selected. The weights are determined by iterating through the preset weight combinations using the same verification samples. If there are no historical samples, the standard item maintenance personnel can first confirm the correspondence between a batch of reported loss items and standard items as initial calibration samples.

[0024] The standard item data is sorted from highest to lowest according to semantic matching scores. Standard items that reach a candidate threshold are designated as candidate standard items. If the number of standard items reaching the candidate threshold is less than the preset number of candidates, higher-ranked standard items are added until the preset number of candidates is reached. Candidate standard item data includes the damaged item identifier, candidate standard item identifier, semantic matching score, candidate ranking, and the score of each matching item. Multiple candidate standard items are retained here; the standard item with the highest semantic matching score is not directly selected as the target standard item.

[0025] S3. Perform multi-source probing queries based on data source field mapping; In this embodiment, step S3 includes the following process: Using each candidate standard item as the query object, the corresponding query field mapping relationship of each claims data source is invoked. For data sources that can directly retrieve standard item identifiers, the candidate standard item identifier is prioritized and written into the code query field; for data sources that only support name queries, the standard item name and standard synonym expression are written into the name query field; then, the fields supported by the data source among the property damage object attributes, handling action attributes, damage reporting scenario attributes, region attributes, and time attributes are written into the query constraint field. This generates trial query data with damage reporting item identifier, candidate standard item identifier, data source identifier, and query round identifier.

[0026] For claims data sources that are available and whose data coverage includes candidate standard item categories, exploratory queries are performed. For claims data sources that are unavailable or whose data coverage does not include candidate standard item categories, no business query requests are sent. Instead, a neutral response is generated with candidate standard item identifiers, data source identifiers, and a neutral reason. This ensures that all configured claims data sources have aggregated feedback records and avoids misinterpreting unexecuted queries as valid misses.

[0027] Different claims data sources use their own trial query templates. For example, a self-built claims project database can be queried using standard project identifier, region, and valid time; a historical claims case database can be queried using project description, damaged object category, handling action, case region, and case closure time; and a market inquiry data source can be queried using project name, material specifications, service region, and quotation time. The field names in the above templates are based on the claims data source configuration data, and different data sources are not required to use the same table structure. To ensure comparability between different candidates, when performing trial queries for the same claim project against different candidates, the damaged object, handling action, and scenario constraints should remain consistent, and only query values ​​related to the candidate standard project should be replaced.

[0028] Probing queries are used to verify whether candidate standard items correspond to actual records, and are not intended for downloading all business data at once. Database-based data sources can use read-only queries with a list of returned fields and a maximum record count; API-based data sources can limit the number of returned records through pagination parameters. Returned fields must include at least the source record identifier, item description, item category or level, object or specification description, and record time; the amount field is only read when it is necessary to determine the item structure or perform subsequent amount allocation. By limiting the number of returned fields and records, the amount of irrelevant records transmitted during the candidate verification phase can be reduced.

[0029] After executing a trial query, the query response is categorized into three execution states: query hit, valid miss, and query exception. A query hit indicates a successful query that returns at least one record; a valid miss indicates a successful query but returns zero records; and a query exception indicates that a request was sent but encountered a statement error, interface error, or a return structure that does not conform to the configuration. Neutral feedback is applied to data sources for which the query was not executed. The execution state, total number of hit records, restricted hit records, query time, neutral feedback, candidate standard item identifier, data source identifier, and query round identifier are correlated to obtain multi-source query feedback data.

[0030] S4. Filter valid query feedback and generate multi-source feedback support data; In this embodiment, step S4 includes the following process: First, filter valid query feedback data based on data source availability, data coverage, and query execution status. Feedback is considered valid when the data source is available, its data coverage includes the candidate standard item category, and the query executes successfully. Feedback indicating query errors, data source unavailability, or data coverage explicitly not including the category is considered neutral and does not reduce the support for candidate standard items. Feedback indicating a successful query but zero records hit is considered a valid miss, indicating that even with query capability and coverage of the corresponding category, the current candidate and constraint have no corresponding records, thus serving as evidence against the candidate.

[0031] The returned records from the query are filtered for validity. First, duplicates are removed based on the source record identifier. Then, the existence of validation fields such as project description, project category or level, object or specification description, and record time is checked to determine field completeness. Records whose field completeness does not meet the preset completeness criteria are not considered as evidence to support or oppose the candidate, but their missing field identifiers are retained for subsequent determination of whether the query fields need to be adjusted. The preset completeness criteria are determined based on the returned fields configured for each claims data source to determine the completeness of the matched records; different claims data sources are not required to have completely identical returned fields.

[0032] For records that meet the field integrity requirements, the item description, object or specification description, and disposal action in the record are matched with the property insurance claim data to obtain record content consistency. Record content consistency can be achieved using the same literal and semantic vector processing method as in step S2, but the comparison object becomes the returned record and the current claim item, and the damaged object and disposal action must be treated as independent comparison fields. This excludes irrelevant records with similar names but inconsistent repair objects or disposal actions. For example, if the claim item requires "replacement," but the returned record only corresponds to "cleaning," even if the item names are similar, a high degree of record content consistency cannot be achieved.

[0033] Project granularity consistency is achieved based on the relationship between the project hierarchy of the returned record and the structure of the reported loss project. When the reported loss project is an independent component and the returned record also corresponds to a single component, their granularity is consistent. When the reported loss project is a comprehensive project, and the returned record can cover its constituent sub-items through a standard hierarchical relationship, granularity consistency is also achieved. However, when the reported loss project is a specific sub-item and the returned record is only an indivisible higher-level comprehensive service, or when the returned record omits necessary components of the reported loss project, project granularity consistency decreases. This allows for the distinction between cases where "names are similar in meaning" and cases where "record structures actually correspond."

[0034] As an optional quantitative implementation method, for the first The candidate standard project in the 1st The hit feedback from each claims data source is used to generate source-level feedback values ​​according to the following formula: ; In the formula, For the first The first claims data source for the first The source-level feedback values ​​of each candidate standard item take values ​​within a closed interval. ; and These are the average record content consistency and the average project granularity consistency of valid hit records that meet the field integrity requirements in the data source, respectively, both of which are within a closed interval. Dimensionless numbers within; The weight used to record content consistency takes a value within a closed interval. Within the project, the weight for consistency in project granularity is: If the query is successful and the total number of records matched is zero, then... Set directly to If the query is successful and there are matching records, but no records meet the field integrity requirements, then... Set as The source is then marked as neutral to avoid calculating the average value for an empty set of records.

[0035] Set effective feedback threshold When the first The claims data source is available, and the data coverage includes the [number]. When the categories of candidate standard items are defined and the query is executed successfully. When the data source is unavailable, the query execution fails, or the data coverage does not include the category, A successful query but no match result still satisfies the requirement. Thus, it participates in the aggregation as opposing evidence. Based on the effective feedback gating value, the aggregated feedback value of the candidate standard item is generated according to the following formula: ; In the formula, For the first Aggregate feedback values ​​of candidate standard items; For the first The reliability weight of each claims data source is a non-negative dimensionless number. The number of configured claims data sources. When the denominator... When the value is zero, no calculation is performed. The candidate standard project was marked as having insufficient feedback to avoid misjudging the lack of effective sources as negative feedback.

[0036] The source support status is determined based on field completeness, record content consistency, and project granularity consistency. When a valid query response from a claims data source contains a matching record that meets both the preset completeness and preset record consistency conditions, the data source is marked as supporting the candidate standard item. When the response is valid but no matching, or when none of the matching records meeting the preset completeness conditions meet the preset record consistency conditions, it is marked as opposing. When the response is marked as neutral, or when none of the matching records meet the preset completeness conditions, it is marked as neutral. Using the above quantitative implementation, source-level support thresholds and source-level opposition thresholds can be set separately, and... To assist in determining the source support status, the preset content consistency threshold, preset granularity consistency threshold, source-level support threshold, source-level opposition threshold, and preset integrity condition are calibrated using the same historical positive and negative sample enumeration method as in step S2, and are saved separately according to the data source. and This is determined by reducing the number of misjudgments in the source support status of historically confirmed samples.

[0037] This section summarizes the source support status of the same candidate standard item across different claims data sources. A cross-source support level is defined as follows: when at least two valid data sources support the same candidate and there are no opposing states meeting the preset conflict conditions; a single-source support level is defined as follows: when there are no supporting states and at least one opposing state, it is defined as follows: when different valid data sources support different candidate standard items, or when the same candidate simultaneously has both supporting and opposing states meeting the preset conflict conditions, it is defined as follows: when there is no valid query feedback data or all claims data sources are in a neutral state, it is defined as follows: insufficient feedback level. Used for comparing feedback strength and identifying conflicts within the same support level, but not for offsetting opposing feedback with semantic matching scores; the target standard item still follows step S6 to first compare the support levels of multi-source feedback, and then sort them within the same support level. Multi-source feedback support data includes candidate standard item identifiers, source support status of each data source, valid support record identifiers, reasons for opposition, support levels, conflict candidate identifiers, and aggregated feedback values.

[0038] The preset conflict conditions are determined based on the number of objection states and the reliability weight of the corresponding claims data source. When the number of objection states reaches a preset threshold for the number of objection sources, or when the reliability weight of the claims data source corresponding to any objection state reaches a preset high-reliability source threshold, the corresponding objection state is determined to have met the preset conflict conditions. The preset threshold for the number of objection sources and the preset high-reliability source threshold are statistically analyzed using historical confirmed semantic mapping records, with the goal of reducing the number of samples where confirmed target standard items are incorrectly judged as source conflicts, and are written into the feedback judgment configuration file.

[0039] S5. When the correction conditions are met, perform targeted correction queries and update feedback support; In this embodiment, step S5 includes the following process: The preset feedback correction conditions include at least one of the following: Candidate standard items with high semantic matching scores are at the unsupported level; all candidate standard items generate valid no-hit feedback in data sources reaching the preset minimum number of available sources; hit records exist, but the consistency of record content or item granularity does not reach the corresponding threshold; different claims data sources support different candidate standard items, resulting in a source conflict level. Query correction data is generated only for candidate standard items and related data sources that meet the feedback correction conditions; repeated queries are not performed on candidates that have already obtained stable cross-source support.

[0040] For valid no-hit responses, first determine if the query field of the data source supports standard item identifiers. If not, read the standard synonyms of candidate standard items and the confirmed data source query aliases to form query term expansion data. If the candidate standard item belongs to a higher-level comprehensive item, read the lower-level standard item names along the standard hierarchy and form query terms accordingly. At least one of the standard synonyms, data source query aliases, and lower-level standard item names is selected based on the existence of the corresponding data, enabling the trial query to verify whether the data source uses aliases or sub-item granularity to store records. During expansion, the attributes of the damaged object and the disposal action attributes are preserved, and completely removing scenario constraints is not the first choice for correction to avoid expanding to irrelevant records.

[0041] For queries that hit but show low consistency in record content, compare each returned field with the reported item data to pinpoint the inconsistencies. If the returned record objects are consistent but the actions taken are inconsistent, add the action to the mandatory query constraint. If the item description field contains a data source-specific abbreviation, extract the common item description from records with high consistency to generate a query alias specific to that data source. If the configured query fields consistently return records lacking item descriptions, switch to an alternative field in the data source configuration that can represent the item's meaning. This results in query field correction data or query constraint supplementary data.

[0042] For source conflict feedback, project descriptions, object descriptions, handling actions, and project levels are extracted from valid records of each supporting source. Descriptions consistent with the reported damage project and capable of distinguishing conflict candidates are retained. These are combined into query terms for conflict verification and sent to available data sources that have not yet provided valid support for the conflict candidates. If the conflict sources differ in project granularity, restricted queries are performed on both higher-level and lower-level candidates. The appropriate approach—whether to use a single target project or a set of target standard projects—is determined based on whether the returned records can cover the structure of the reported damage project.

[0043] The trial query is executed again based on the corrected query data. This second query only targets the marked candidate criteria items and data sources, and continues to apply the return field limitations and the maximum number of query records. The initial feedback and corrected feedback are deduplicated according to the data source identifier, source record identifier, and candidate criteria item identifier. For the execution status of the same query condition, the status of the last successful execution is used, while retaining the original query rounds and correction types as query evidence. Subsequently, the source support status and multi-source feedback support level are recalculated according to step S4. If there is still no effective support after the query correction rounds reach the preset limit, automatic expansion is stopped and a pending review flag is generated to prevent unlimited relaxation of conditions from causing a full-scale search.

[0044] S6. Determine the target standard items and generate semantic alignment results; In this embodiment, step S6 includes the following process: Candidate standard items without a support level are removed from the candidate standard item data. For the remaining candidates, the support levels of multi-source feedback are compared first, with cross-source support levels taking precedence over single-source support levels. If the support levels are the same, the consistency of the record content of the effective support records, the consistency of the item granularity, and the semantic matching score are compared in that order. By comparing feedback support first and then using semantic matching scores for peer sorting, it is possible to avoid selecting candidate standard items that do not have corresponding records in the actual claims data source simply because the standard name is closer.

[0045] When a candidate standard item ranked first can independently cover the structure of a reported loss item, it is identified as the target standard item and a one-to-one mapping relationship is established. When a reported loss item is a comprehensive item, and trial query feedback indicates that different subordinate standard items correspond to its components, multiple subordinate standard items are identified as the target standard item set and a one-to-many mapping relationship is established. When multiple reported loss item data points correspond to the same standard item in both the standard level and the effective support records, a many-to-one mapping relationship can be established. The semantic alignment result includes at least the reported loss item identifier, the target standard item identifier, the mapping relationship type, the semantic matching score, the multi-source feedback support level, the effective support record identifier, and the alignment status.

[0046] When a standardized loss report list needs to be generated, for a one-to-one mapping relationship, the original loss amount of the corresponding property insurance loss report data is directly used as the loss amount of the target standard project; for a one-to-many mapping relationship, the original loss amount is allocated according to the proportion of the reference amount in the valid supporting records for each target standard project. If the reference amount is missing, an initial allocation result to be reviewed is generated according to the preset allocation rules; for a many-to-one mapping relationship, the original loss amounts of the corresponding property insurance loss report data are merged. After the amount allocation or merging is completed for one-to-many or many-to-one mapping relationships, the rounding difference caused by the smallest unit of currency is added to the preset rounding difference receiving project to keep the total loss amount consistent before and after semantic alignment. The amount processing is used to maintain the consistency of the list structure and does not participate in the semantic support judgment of the target standard project.

[0047] When the updated multi-source feedback support data is still at the source conflict level or insufficient feedback level, or when there are no candidate standard items that meet the minimum support level, a unique target standard item is not forced to be output. Instead, a semantic alignment result to be reviewed is generated. The semantic alignment result to be reviewed retains the candidate ranking, the execution status of each data source, the reasons for support or opposition, the corrected query content, and the summary of valid records for the loss assessment personnel to confirm. After the confirmation result is returned, the final target standard item is then written into the semantic alignment result.

[0048] After the target standard project is reviewed and confirmed or confirmed through subsequent processes, a confirmed semantic mapping record is established between the reported project name, damaged object attributes, and handling action attributes and the target standard project identifier. Project descriptions and valid query fields used by each claims data source are extracted from the valid query feedback data supporting the target standard project to form a data source query alias record. The confirmed semantic mapping record is used for the generation of subsequent candidate standard projects, and the data source query alias record is used for subsequent trial query conversion. Unconfirmed results are not written to the semantic alignment mapping library to avoid the accumulation of erroneous feedback.

[0049] For example, for the reported item "complete restoration of the bathroom," the candidate in the standard item library closest in meaning to its text might be the higher-level comprehensive renovation item. However, the historical claims database and market inquiry data source might store records separately for sub-items such as wall finishes, floor finishes, sanitary ware, and ceilings. By performing a trial query on the higher-level candidate, effective no-match or lower-level item consistency can be obtained. Based on this, the processor generates lower-level item queries along the standard hierarchy and obtains records consistent with the reported item and disposal action from multiple data sources, ultimately establishing a one-to-many mapping between the reported item and multiple target standard items. This example is only for illustrating the feedback correction relationship; the specific standard item names and hierarchy are subject to the adopted standard item library.

[0050] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A semantic alignment method for property insurance loss reporting data based on multi-source query feedback, characterized in that, include: Acquire property insurance loss reporting data, standard project data, and claims data source configuration data. The claims data source configuration data includes query field mapping relationships, data coverage, and data source availability status. Semantic matching is performed between property insurance claim data and standard data, and semantic retrieval is performed on the standard data based on the semantic matching results to obtain candidate standard data containing semantic matching scores. Based on the query field mapping relationship, the data of each candidate standard item is converted into trial query data. Trial queries are performed on the available claims data sources whose data coverage includes the candidate standard item categories. Neutral feedback is generated for the unavailable claims data sources or those that do not cover the candidate standard item categories, resulting in multi-source query feedback data. Based on query execution status, number of hit records, consistency of record content, and consistency of project granularity, the feedback from each claims data source to the candidate standard project is divided into supportive, opposing, or neutral status, and multi-source feedback support data is generated. When the multi-source feedback support data meets the preset feedback correction conditions, query correction data is generated based on the multi-source query feedback data, and the trial query is executed again to update the multi-source feedback support data. Based on semantic matching scores and multi-source feedback support data, target standard items are determined, and semantic alignment results between property insurance loss reporting items and target standard items are generated.

2. The semantic alignment method for property insurance loss reporting data based on multi-source query feedback as described in claim 1, characterized in that, Property insurance loss reporting data, standard item data, and claims data source configuration data specifically include: Property insurance loss reporting data includes loss reporting item identifier, loss reporting item name, damaged object attributes, disposal action attributes, loss reporting scenario attributes, item structure relationship, and original loss reporting amount; Standard project data includes standard project identifier, standard project name, standard synonyms, standard hierarchy, and standard project category; The claims data source configuration data also includes data source identifier, access interface, query fields, return fields, and maximum number of query records. Each claims data source includes at least two of the following: self-built claims project database, historical claims case database, and market inquiry data source.

3. The semantic alignment method for property insurance loss reporting data based on multi-source query feedback as described in claim 1, characterized in that, The candidate standard item data containing semantic matching scores is obtained, specifically including: Extract the name of the reported loss item, the attributes of the damaged object, the attributes of the disposal action, and the attributes of the reported loss scenario from the property insurance loss item data to generate semantic feature data of the loss item; The semantic feature data of the reported damage items are calculated separately for the literal matching degree, semantic vector matching degree, scene matching degree, and standard level matching degree between the semantic feature data of each standard item and the data of each standard item, and a semantic matching score is generated accordingly. The standard project data is sorted according to the semantic matching score, and the standard project data that meets the preset candidate conditions and its semantic matching score are associated as candidate standard project data.

4. The semantic alignment method for property insurance loss reporting data based on multi-source query feedback as described in claim 3, characterized in that, We obtained multi-source query feedback data, specifically including: Based on the query field mapping relationship, the standard project identifier, standard project name or standard synonym of the candidate standard project is written into the query field corresponding to each claim data source, and the property damage object attribute, handling action attribute and loss reporting scenario attribute is written into the corresponding query constraint field to generate trial query data with candidate standard project identifier and data source identifier; Send probing query data to a claims data source whose data source availability status is available and whose data coverage includes candidate standard item categories, so that it returns the query execution status, the total number of hit records, and the number of hit records not exceeding the upper limit of the number of query records; Generate neutral feedback for claims data sources that are unavailable or whose data coverage does not include candidate standard item categories; The query execution status, total number of hit records, number of hit records, and neutral feedback are associated with the candidate standard item identifier and data source identifier to form multi-source query feedback data.

5. A semantic alignment method for property insurance loss reporting data based on multi-source query feedback as described in claim 1, characterized in that, Generate multi-source feedback support data for each candidate standard item, specifically including: When the claims data source is available, the data coverage includes candidate standard item categories, and the query execution status indicates that the query was successful, the corresponding multi-source query feedback data will be determined as valid query feedback data. Valid query feedback data that was successfully queried but had a total of zero records hits was marked as valid no-hit feedback, and multi-source query feedback data that was abnormal, had an unavailable data source, or whose data coverage did not include candidate standard item categories was marked as neutral feedback. The field integrity of the hit records is judged. The project description and project level in the hit records that meet the preset integrity conditions are matched with the property insurance loss report project data for content and structure, respectively, to obtain the consistency of record content and project granularity. Based on valid no-hit feedback, field completeness, number of hit records, record content consistency, and project granularity consistency, the source support status of each claims data source for the candidate standard project is determined, and the data is aggregated into multi-source feedback support data according to the data source identifier.

6. A semantic alignment method for property insurance loss reporting data based on multi-source query feedback as described in claim 5, characterized in that, The source support status includes support, opposition, and neutral status, which are aggregated into multi-source feedback support data according to the data source identifier, specifically including: When there are matching records in the valid query feedback data that meet the preset completeness condition and the preset record consistency condition, a support status is generated; When the valid query feedback data is a valid no-hit feedback, or when the hit records that meet the preset complete condition do not meet the preset record consistency condition, an opposition status is generated; A neutral state is generated when multi-source query feedback data is marked as neutral feedback, or when the number of hit records does not meet the preset completeness condition. When at least two valid data sources support the same candidate standard item and there is no objection state that meets the preset conflict conditions, a cross-source support level is determined; when only one valid data source supports the candidate standard item and there is no objection state that meets the preset conflict conditions, a single-source support level is determined; when there is no support state and at least one objection state, a no-support level is determined; when different valid data sources support different candidate standard items, or when the same candidate standard item has both support and objection states that meet the preset conflict conditions, a source conflict level is determined; when there is no valid query feedback data or all claim data sources are in a neutral state, an insufficient feedback level is determined.

7. A semantic alignment method for property insurance loss reporting data based on multi-source query feedback as described in claim 1, characterized in that, Query correction data is generated based on multi-source query feedback data, specifically including: When a candidate standard item corresponds to a valid no-hit response, at least one of the lower-level standard items or standard synonyms of the candidate standard item is read based on the standard hierarchy relationship and standard synonym expression data to generate query term extended data. When the consistency of the content of the hit record does not reach the preset content consistency threshold, the returned fields that are inconsistent with the property insurance loss report data are determined from the hit record, and query field correction data or query constraint supplementary data are generated. When multi-source feedback supports data indicating source conflict, extract the project description, object description, or handling action that is consistent with the property insurance loss report data from the hit records in the support state, and generate query terms to be verified against the conflict. Associate query term extension data, query field correction data, query constraint supplementary data, or query terms to be verified for conflict with the corresponding candidate standard item identifier and data source identifier to form query correction data.

8. A semantic alignment method for property insurance loss reporting data based on multi-source query feedback as described in claim 7, characterized in that, Based on semantic matching scores and updated multi-source feedback support data, target criteria items are determined, specifically including: Remove candidate standard projects that are at the "no support" level from the candidate standard project data; Candidate standard projects are selected in order of cross-source support level taking precedence over single-source support level. Within the same support level, they are ranked in order of record content consistency, project granularity consistency, and semantic matching score. The candidate standard projects ranked first are determined as the target standard projects. When the updated multi-source feedback support data is still at the source conflict level or insufficient feedback level, or there are no candidate standard items that meet the preset minimum support level, a semantic alignment result to be reviewed is generated.

9. A semantic alignment method for property insurance loss reporting data based on multi-source query feedback as described in claim 1, characterized in that, The target standard items include a single target standard item or a set of target standard items, generating semantic alignment results between property insurance claim data and target standard items, specifically including: Based on the project structure relationship and the standard hierarchy relationship of the target standard project in the property insurance loss reporting project data, determine the one-to-one mapping relationship, one-to-many mapping relationship or many-to-one mapping relationship; When a single target standard item can independently correspond to the item structure of a single property insurance claim item data, a one-to-one mapping relationship is determined, and the original claim amount of the property insurance claim item data is used as the claim amount of the corresponding target standard item. When a one-to-many mapping relationship is determined, the original loss amount is allocated according to the proportion of the reference amount in the corresponding hit record of each target standard item; when the reference amount is missing, the initial allocation result to be reviewed is generated according to the preset allocation rules; when a many-to-one mapping relationship is determined, the original loss amount of the corresponding property insurance loss item data is merged. Perform tail difference correction on the amount allocation or merging results in one-to-many or many-to-one mapping relationships to keep the total amount of loss consistent before and after semantic alignment, and associate the mapping relationship, target standard item identifier and corresponding loss amount as semantic alignment result.

10. A semantic alignment method for property insurance loss reporting data based on multi-source query feedback as described in claim 8, characterized in that, After generating the semantic alignment results, the following is also included: After the target standard items are reviewed and confirmed, a confirmed semantic mapping record will be established between the name of the reported loss item, the attribute of the damaged object, the attribute of the disposal action in the property insurance loss reporting item data and the target standard item identifier; Extract the project description and valid query fields used by each claims data source from the valid query feedback data that supports the target standard project, and generate data source query alias records; Confirmed semantic mapping records and data source query alias records are written into the semantic alignment mapping library for subsequent generation of candidate standard items for property insurance loss reporting data and conversion of trial query data. Semantic alignment results that have not been reviewed and confirmed are not written into the semantic alignment mapping library.