Insurance parameter configuration method and system based on multi-modal data and knowledge graph

CN122820344APending Publication Date: 2026-09-25BAOTENG NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610895564.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-22
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

一方面,多模态数据的处理方式割裂了不同模态之间的互补性与协同性,例如投保人的面部表情、病史文本与体征数据本可共同反映真实健康状况,但现有方法仅能独立提取各模态的特征,难以建立跨模态的一致表征,导致不完整信息下的风险评估偏差

Benefits of technology

[0016]本发明利用多模态数据与知识图谱的深度融合,显著提升了投保参数配置的精准度与全面性。多模态数据涵盖结构化与非结构化特征,通过跨模态对齐处理后形成的统一表征向量,能够弥合不同数据源之间的语义鸿沟,避免单一模态信息不完整或偏差导致的评估失真。知识图谱的构建将投保对象、属性及其关联关系进行结构化表达,基于多跳关系路径传播可自动挖掘隐含的风险传导链条,发现直接观测难以捕获的间接风险因素,使风险评估覆盖范围更广、层次更深。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820344A_ABST
    Figure CN122820344A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of data processing, and more particularly to a kind of based on multi-modal data and knowledge graph's insurance parameter configuration method and system.The method includes obtaining the multi-modal data of the object to be insured and extracting structured and unstructured feature vectors, obtaining uniform representation vector by cross-modal alignment of the two, constructing knowledge graph based on uniform representation vector and domain ontology rules and determining associated entities and attributes through multi-hop relationship path, determining insurance parameter configuration scheme by constraint reasoning accordingly, and adjusting cross-modal alignment parameters using execution feedback data.The present application improves the accuracy and adaptability of insurance parameter configuration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method and system for configuring insurance parameters based on multimodal data and knowledge graphs. Background Technology

[0002] In the field of insurance underwriting parameter configuration technology, current conventional practices mainly rely on single-modal data sources, such as spreadsheets or questionnaires submitted by policyholders, combined with statistical models or pre-set rule engines for risk assessment and premium calculation. Some systems attempt to introduce unstructured data such as images or voice, but these are usually processed independently, converted into single features, and then simply concatenated with structured data, lacking in-depth mining of the inherent relationships between multimodal data. The application of knowledge graphs is still in its early stages, mostly used to store static relationships between insurance products and terms, without effectively linking them to the dynamic characteristics of the insured.

[0003] This conventional approach has at least two significant drawbacks. Firstly, the processing of multimodal data fragments the complementarity and synergy between different modalities. For example, an insured's facial expressions, medical history, and vital signs could collectively reflect their true health status, but current methods can only extract features from each modality independently, making it difficult to establish a consistent cross-modal representation. This leads to biased risk assessments based on incomplete information. Secondly, the use of knowledge graphs remains at the level of shallow entity queries and simple rule matching, lacking the ability for path-based reasoning and failing to dynamically discover indirect connections between the insured and risk factors (such as long-term correlations between past medical history and specific diseases). This makes insurance parameter configuration schemes often too rigid and difficult to adapt to personalized risk scenarios.

[0004] The aforementioned shortcomings result in insufficient accuracy of existing technologies in complex insurance scenarios, and parameter adjustments rely on human experience, making automated optimization impossible. Therefore, effectively integrating multimodal data and constructing a reasonable knowledge support system has become a pressing technical problem to be solved in this field. Summary of the Invention

[0005] This invention provides a method and system for configuring insurance parameters based on multimodal data and knowledge graphs, which can solve the problems in the prior art.

[0006] A first aspect of this invention provides a method for configuring insurance application parameters based on multimodal data and knowledge graphs, comprising: Obtain multimodal data of the insured object, and extract features from the multimodal data to obtain structured feature vectors and unstructured feature vectors; The structured feature vector and the unstructured feature vector are aligned across modes to obtain a unified representation vector. Based on the unified representation vector and the preset domain ontology rules, a knowledge graph containing entity nodes, attribute nodes, and relation edges of the insured object is determined. Starting from the entity node corresponding to the insured object in the knowledge graph, the set of associated entities and the set of associated attributes related to risk assessment are determined through multi-hop relation path propagation. Based on the set of associated entities, the set of associated attributes, and the unified representation vector, an insurance parameter configuration scheme is determined through constraint reasoning. Execution feedback data of the insurance parameter configuration scheme is obtained, and the parameters in the cross-modal alignment process are adaptively adjusted using the execution feedback data.

[0007] Obtain multimodal data of the insured object, and perform feature extraction on the multimodal data to obtain structured feature vectors and unstructured feature vectors, including: The structured attribute data and unstructured descriptive data of the insured object are obtained as the multimodal data, wherein the structured attribute data is represented by key-value pairs and the unstructured descriptive data is represented by text sequences; The structured attribute data is subjected to field standardization and missing value imputation to obtain standardized structured data; the unstructured description data is subjected to word segmentation and stop word filtering to obtain standardized unstructured data. The numerical fields in the standardized structured data are mapped to a unified numerical space through normalization transformation, and the categorical fields in the standardized structured data are converted into vector representations through encoding mapping. The normalized result of the numerical field is concatenated with the vector representation of the categorical field to obtain the structured feature vector; The normalized unstructured data is input into a pre-trained semantic representation network to extract semantic feature representations from the normalized unstructured data. The semantic feature representations are then pooled and aggregated to obtain the unstructured feature vector.

[0008] The structured feature vector and the unstructured feature vector are subjected to cross-modal alignment to obtain a unified representation vector, including: Semantic space projection processing is performed on the structured feature vector and the unstructured feature vector respectively, mapping the structured feature vector to obtain a first mapping vector and the unstructured feature vector to obtain a second mapping vector; Calculate the similarity distribution between the first mapping vector and the second mapping vector at multiple semantic granularity levels to determine a multi-granularity cross-modal correspondence matrix. The multi-granularity cross-modal correspondence matrix simultaneously characterizes the correlation strength between the structured feature vector and the unstructured feature vector at both the global and local semantic levels. Based on the multi-granularity cross-modal correspondence matrix, fusion weights are assigned to each dimension component of the first mapping vector and the second mapping vector, and the fusion weights are constrained and adjusted according to the risk type attribute of the insured object; The first mapping vector and the second mapping vector are weighted and combined according to the fusion weights adjusted by the constraints to obtain the unified representation vector.

[0009] Based on the unified representation vector and preset domain ontology rules, a knowledge graph containing entity nodes, attribute nodes, and relation edges of the insured object is determined. Starting from the entity node corresponding to the insured object, the knowledge graph is propagated through multi-hop relation paths to determine the set of associated entities and the set of associated attributes related to risk assessment, including: Based on the unified representation vector, the entity type of the object to be insured in the preset domain ontology rules is identified by semantic matching, and the entity node of the object to be insured is generated. Based on the feature dimension distribution in the unified representation vector, the attribute information of the object to be insured is extracted through semantic parsing, and the attribute information is labeled with type and hierarchically divided according to the attribute definition in the domain ontology rules to generate an attribute node set; Based on the relation type constraints defined in the domain ontology rules, relation edges are established between the entity nodes and the attribute node set to obtain the knowledge graph; Starting from the entity node corresponding to the insured object, a multi-hop traversal is performed along the relation edges. During the traversal, the propagation path is filtered according to the semantic type of the relation edges to obtain a set of candidate association paths. Extract entity nodes related to risk assessment semantics from the termination nodes of the candidate association path set to obtain the association entity set; The associated attribute set is obtained by extracting the attribute nodes related to risk quantification calculation from the attribute nodes associated with each entity node in the associated entity set.

[0010] Based on the feature dimension distribution in the unified representation vector, the attribute information of the insured object is extracted through semantic parsing. The attribute information is then labeled with types and hierarchically divided according to the attribute definitions in the domain ontology rules, generating a set of attribute nodes, including: Semantic parsing is performed on the feature dimension distribution in the unified representation vector to identify the semantic concepts corresponding to each feature dimension, thereby obtaining a set of semantic concepts. The semantic concepts in the semantic concept set are semantically matched with the predefined attribute definitions in the domain ontology rules to determine the mapping relationship between semantic concepts and attribute types, and the corresponding attribute information is extracted for the insured object based on the mapping relationship. Based on the attribute inheritance and attribute dependency relationships defined in the domain ontology rules, hierarchical analysis is performed on the attribute information to determine the hierarchical relationships and dependency constraints between each attribute information. The attribute information is organized into a tree-like hierarchical structure according to the hierarchical relationship, and constraint markers are determined between the nodes of the tree-like hierarchical structure based on the dependency constraint relationship. The attribute information of each level in the tree hierarchy and the constraint markers are instantiated as attribute nodes to obtain the attribute node set.

[0011] Based on the set of associated entities, the set of associated attributes, and the unified representation vector, an insurance parameter configuration scheme is determined through constraint reasoning. Execution feedback data for the insurance parameter configuration scheme is obtained, and the parameters in the cross-modal alignment process are adaptively adjusted using this execution feedback data, including: Knowledge graph logic rules are extracted from the set of associated entities and the set of associated attributes, and the semantic similarity between each associated entity in the set of associated entities and the insured object is calculated based on the unified representation vector to obtain the semantic similarity distribution; Based on the logical rules of the knowledge graph, the value space of candidate insurance parameters is constrained and filtered to obtain a set of compliant parameter values. Then, the combination of parameter values ​​in the set of compliant parameter values ​​is prioritized according to the semantic similarity distribution to obtain the insurance parameter configuration scheme. Obtain execution feedback data of the insurance parameter configuration scheme in actual insurance business, and calculate the deviation between the insurance parameter configuration scheme and actual business needs based on the execution feedback data; The feature vector causing the deviation in the multimodal data is determined based on the deviation amount, and the parameters in the cross-modal alignment process are adaptively adjusted based on the feature vector causing the deviation.

[0012] Based on the logical rules of the knowledge graph, the value space of candidate insurance parameters is constrained and filtered to obtain a set of compliant parameter values. Then, according to the semantic similarity distribution, the combinations of parameter values ​​in the compliant parameter value set are prioritized to obtain the insurance parameter configuration scheme, including: The logical rules of the knowledge graph are parsed to obtain a set of parameter boundary conditions and a set of parameter consistency conditions. The set of parameter boundary conditions limits the upper and lower bounds of the value of a single parameter, and the set of parameter consistency conditions limits the value association relationship between multiple parameters. The compliance of each parameter value combination in the candidate insurance parameter value space is verified. It is determined whether the parameter value combination simultaneously satisfies the parameter boundary condition set and the parameter consistency condition set. The parameter value combinations that satisfy all conditions are retained and the parameter value combinations that violate any condition are removed to obtain the compliant parameter value set. For each combination of parameter values ​​in the set of compliance parameter values, a comprehensive semantic similarity score is calculated between the associated entity corresponding to the parameter value combination and the insured object based on the semantic similarity distribution. The parameter value combinations in the compliance parameter value set are sorted in descending order according to the comprehensive semantic similarity score, and the parameter value combination with the highest sorted position is selected as the insurance parameter configuration scheme.

[0013] A second aspect of this invention provides an insurance parameter configuration system based on multimodal data and knowledge graphs, comprising: The feature extraction unit is used to acquire multimodal data of the insured object, extract features from the multimodal data, and obtain structured feature vectors and unstructured feature vectors. The alignment processing unit is used to perform cross-modal alignment processing on the structured feature vector and the unstructured feature vector to obtain a unified representation vector. The graph construction unit is used to determine a knowledge graph containing entity nodes, attribute nodes and relation edges of the insured object based on the unified representation vector and the preset domain ontology rules. In the knowledge graph, starting from the entity node corresponding to the insured object, the set of associated entities and the set of associated attributes related to risk assessment are determined through multi-hop relation path propagation. The configuration adjustment unit is used to determine the insurance parameter configuration scheme through constraint reasoning based on the associated entity set, the associated attribute set, and the unified representation vector, obtain the execution feedback data of the insurance parameter configuration scheme, and use the execution feedback data to adaptively adjust the parameters in the cross-modal alignment process.

[0014] A third aspect of the present invention provides an electronic device, comprising: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0015] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0016] This invention leverages the deep fusion of multimodal data and knowledge graphs to significantly improve the accuracy and comprehensiveness of insurance parameter configuration. Multimodal data encompasses both structured and unstructured features. The unified representation vector formed through cross-modal alignment can bridge the semantic gap between different data sources, avoiding assessment distortion caused by incomplete or biased information from a single modality. The construction of the knowledge graph provides a structured representation of insured objects, attributes, and their relationships. Based on multi-hop relationship path propagation, it can automatically uncover hidden risk transmission chains and discover indirect risk factors that are difficult to capture through direct observation, thus broadening the scope and deepening the level of risk assessment.

[0017] This invention utilizes a constraint reasoning mechanism driven by unified representation vectors and domain ontology rules to automatically generate insurance parameter configuration schemes that comply with regulatory requirements and business logic. This process eliminates the need for manually pre-setting fixed rule templates; instead, it dynamically matches the strength of association between entities and attributes to achieve differentiated parameter recommendations. Compared to traditional manual configuration or simple rule engines, this invention reduces the risk of subjective bias and rule omissions, while significantly shortening the configuration cycle and improving business processing efficiency.

[0018] This invention employs an adaptive adjustment strategy for cross-modal alignment parameters, relying on continuous driving forces from execution feedback data to enable the model to correct feature alignment methods in real time based on actual claims or operational results. This closed-loop optimization mechanism enhances the system's adaptability to changes in business scenarios. For example, when external environments or the status of insured individuals fluctuate, alignment parameters can automatically converge to the optimal state, thereby maintaining the long-term effectiveness of the configuration scheme. Overall, this invention achieves end-to-end intelligent management from data collection to scheme implementation and continuous optimization, significantly improving the reliability, flexibility, and automation level of insured parameter configuration. Attached Figure Description

[0019] Figure 1 This is a flowchart illustrating the insurance parameter configuration method based on multimodal data and knowledge graphs according to an embodiment of the present invention. Figure 2 This is a schematic diagram illustrating the process of determining a unified representation vector according to an embodiment of the present invention. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0022] Figure 1 This is a flowchart illustrating the insurance parameter configuration method based on multimodal data and knowledge graphs according to an embodiment of the present invention. Figure 1 As shown, the method includes: Obtain multimodal data of the insured object, and extract features from the multimodal data to obtain structured feature vectors and unstructured feature vectors; The structured feature vector and the unstructured feature vector are aligned across modes to obtain a unified representation vector. Based on the unified representation vector and the preset domain ontology rules, a knowledge graph containing entity nodes, attribute nodes, and relation edges of the insured object is determined. Starting from the entity node corresponding to the insured object in the knowledge graph, the set of associated entities and the set of associated attributes related to risk assessment are determined through multi-hop relation path propagation. Based on the set of associated entities, the set of associated attributes, and the unified representation vector, an insurance parameter configuration scheme is determined through constraint reasoning. Execution feedback data of the insurance parameter configuration scheme is obtained, and the parameters in the cross-modal alignment process are adaptively adjusted using the execution feedback data.

[0023] Obtain multimodal data of the insured object, and perform feature extraction on the multimodal data to obtain structured feature vectors and unstructured feature vectors, including: The structured attribute data and unstructured descriptive data of the insured object are obtained as the multimodal data, wherein the structured attribute data is represented by key-value pairs and the unstructured descriptive data is represented by text sequences; The structured attribute data is subjected to field standardization and missing value imputation to obtain standardized structured data; the unstructured description data is subjected to word segmentation and stop word filtering to obtain standardized unstructured data. The numerical fields in the standardized structured data are mapped to a unified numerical space through normalization transformation, and the categorical fields in the standardized structured data are converted into vector representations through encoding mapping. The normalized result of the numerical field is concatenated with the vector representation of the categorical field to obtain the structured feature vector; The normalized unstructured data is input into a pre-trained semantic representation network to extract semantic feature representations from the normalized unstructured data. The semantic feature representations are then pooled and aggregated to obtain the unstructured feature vector.

[0024] The multimodal data of potential insured individuals originates from multiple heterogeneous channels, encompassing structured sources such as insurance application forms, historical claims records, medical and health records, and property appraisal reports, as well as unstructured sources such as risk description texts in natural language, accounts of accidents, and customer questionnaire responses. Structured attribute data is organized in key-value pairs, with each field name corresponding to a specific attribute value, such as "Age: 35," "Occupation Category: High Risk," and "Past Medical History: None." This representation facilitates field-level standardization and statistical analysis. Unstructured descriptive data exists as continuous text sequences, containing a large amount of semantic information expressed in natural language, requiring language models to effectively extract its semantic features.

[0025] For structured attribute data, the first step is field standardization, which unifies inconsistent field naming, units, and formatting across different data sources. For example, different expressions such as "birth date," "birthday," and "DOB" are mapped to the standard field "birth_date," and the date format is converted to the ISO 8601 standard format. After field standardization, missing values ​​are imputed. For numerical missing fields, the mean or median of samples in the same category is used; for categorical missing fields, the most frequent category value or a dedicated "unknown" category label is used. This ensures that missing values ​​do not introduce noise or cause calculation errors during subsequent feature extraction, resulting in standardized structured data.

[0026] For unstructured description data, word segmentation processing and stop word filtering processing are performed to obtain standardized unstructured data. Word segmentation processing cuts continuous text into a word sequence. For Chinese text, a word segmentation method combining a dictionary and a statistical model is adopted, while for English text, segmentation is performed with spaces and punctuation marks as boundaries. Stop word filtering processing removes high-frequency words that contribute extremely little to semantics, such as "de", "le", "shi", "and", "the", from the word sequence, which reduces noise interference in subsequent semantic modeling, reduces the size of the vocabulary, and improves the efficiency and accuracy of semantic feature extraction. After word segmentation and stop word filtering, the unstructured description data is transformed into a standardized sequence composed of effective semantic words, providing a clean text basis for the input of the subsequent semantic representation network.

[0027] Standardized structured data includes two types: numerical fields and categorical fields, and different mapping strategies need to be adopted respectively to convert them into vector representations. For numerical fields, the original values are mapped to a unified numerical space through normalization transformation. Let the original value of a certain numerical field be , the minimum value of this field in the training set is , the maximum value is , then the normalized value is calculated by the following formula: ; Normalization processing eliminates the influence of dimensional differences between different numerical fields on feature fusion, enabling fields with vastly different numerical ranges such as age, income, and housing area to participate in subsequent calculations on the same scale. For categorical fields, discrete categories are converted into dense vector representations through encoding mapping. When the number of categories is small, one-hot encoding is used to map each category to a sparse binary vector; when the number of categories is large, embedding encoding is adopted, which maps each category to a low-dimensional dense vector through a learnable embedding matrix, thereby avoiding the dimension explosion problem. The parameters of embedding encoding are jointly optimized with other parameters during the overall model training, so that categories with similar semantics are closer to each other in the vector space.

[0028] Arrange the normalization results of all numerical fields into a numerical sub-vector, arrange the vector representations of all categorical fields into a categorical sub-vector in sequence, then concatenate the numerical sub-vector and the categorical sub-vector along the feature dimension to obtain a structured feature vector. The concatenation operation directly merges the two parts of features along the vector dimension, retains the independent feature information of each field, and provides a complete structured information basis for subsequent cross-modal alignment processing. The dimension of the structured feature vector is determined by the sum of the number of numerical fields and the encoding dimension of categorical fields, and can be flexibly adjusted according to the specific field settings of the insurance underwriting scenario in practical applications.

[0029] Normalized unstructured data undergoes deep semantic feature extraction using a pre-trained semantic representation network. This network employs a pre-trained language model based on the Transformer architecture, pre-trained on large-scale general corpora and insurance-specific corpora, enabling it to understand professional semantics such as insurance terminology, risk descriptions, and medical expressions. The word sequences of the normalized unstructured data are converted into word sequences and marked with special start tags before being input into the pre-trained semantic representation network. The network models the semantic dependencies between word units through a multi-layer self-attention mechanism, outputting a context-aware semantic feature representation for each word unit position.

[0030] Pooling aggregation is performed on the semantic feature representations at the word level, compressing the positional features of the variable-length sequence into a fixed-dimensional unstructured feature vector. Pooling aggregation can employ a mean pooling strategy, which takes the arithmetic average of the feature vectors at all word positions, ensuring that the semantic information of each word in the sequence contributes equally to the final vector. Alternatively, a max pooling strategy can be used, taking the maximum activation value among all words in each feature dimension, focusing on retaining the most significant semantic signals. In practical insurance parameter configuration scenarios, a weighted pooling strategy can also be used, assigning different weights to words based on their importance in the risk assessment task, so that risk-related keywords contribute more to the final unstructured feature vector. After pooling aggregation, a fixed-dimensional unstructured feature vector is obtained. This vector condenses the core semantic information of the original text description, laying the foundation for subsequent cross-modal alignment processing with the structured feature vector.

[0031] Through the above processing flow, structured attribute data and unstructured descriptive data are converted into structured feature vectors and unstructured feature vectors, respectively. The two types of vectors characterize the risk features of the insured object from different dimensions and together constitute the output of the multimodal feature extraction stage, supporting the complete process of subsequent knowledge graph construction and insurance parameter configuration.

[0032] Figure 2 This is a schematic diagram illustrating the process of determining a unified representation vector according to an embodiment of the present invention. Figure 2 As shown, the structured feature vector and the unstructured feature vector are subjected to cross-modal alignment processing to obtain a unified representation vector, including: Semantic space projection processing is performed on the structured feature vector and the unstructured feature vector respectively, mapping the structured feature vector to obtain a first mapping vector and the unstructured feature vector to obtain a second mapping vector; Calculate the similarity distribution between the first mapping vector and the second mapping vector at multiple semantic granularity levels to determine a multi-granularity cross-modal correspondence matrix. The multi-granularity cross-modal correspondence matrix simultaneously characterizes the correlation strength between the structured feature vector and the unstructured feature vector at both the global and local semantic levels. Based on the multi-granularity cross-modal correspondence matrix, fusion weights are assigned to each dimension component of the first mapping vector and the second mapping vector, and the fusion weights are constrained and adjusted according to the risk type attribute of the insured object; The first mapping vector and the second mapping vector are weighted and combined according to the fusion weights adjusted by the constraints to obtain the unified representation vector.

[0033] After obtaining structured and unstructured feature vectors, it is necessary to map the feature representations from different modalities to the same semantic space to eliminate distribution bias caused by modality differences. For structured feature vectors, a learnable linear projection matrix is ​​used. Transform it and map it to the target semantic space to obtain the first mapping vector. For unstructured feature vectors, another independent linear projection matrix is ​​used. Transform it to obtain the second mapping vector. The parameter dimensions of both projection matrices are aligned with the dimensions of the target semantic space, ensuring... and Having the same vector dimension provides a unified numerical basis for subsequent similarity calculations. During the projection process, nonlinear activation functions (such as ReLU or GELU) can be selectively introduced to enhance the expressiveness of the mapping, enabling the projection results to capture the potential nonlinear semantic structure in the original features.

[0034] After completing the semantic space projection, computation needs to be performed at multiple semantic granularity levels. and The similarity distribution between them, with multiple semantic granularity levels, refers to different abstract scales from overall semantics (global level) to fine-grained semantic components (local level). At the global semantic level, and Treating the whole as two semantic units, calculate the cosine similarity. This is used to measure the degree of consistency between two modal features in the overall semantic direction. At the local semantic level, it will... and The vectors are divided into several sub-vector segments according to a preset segmentation strategy. The cosine similarity is calculated for each pair of corresponding sub-vector segments to obtain a local similarity sequence. ,in The index of the sub-vector segment is used to characterize the fine-grained correspondence between two modal features across different semantic dimension intervals. Global similarity and local similarity sequences are integrated to construct a multi-granularity cross-modal correspondence matrix. Each element of the matrix Indicates the first The semantic granularity level is the first The matrix encodes the correlation strength across the dimensional component intervals. It simultaneously encodes the correspondence between two modal features in both macroscopic overall semantics and microscopic local semantics, avoiding information loss caused by single-granularity similarity calculations.

[0035] Based on multi-granularity cross-modal correspondence matrix ,for and The fusion weights are assigned to the components of each dimension. Specifically, for the matrix... Weighted aggregation is performed along the granularity level dimension to obtain the original weight distribution vector corresponding to each dimensional component. ,in Each element Reflecting the The importance of dimensional components in cross-modal alignment. Dimensional components with higher association strength have larger corresponding fusion weights, indicating stronger semantic consistency between the two modalities and thus requiring greater information contribution during the fusion process. Normalization is performed to ensure that the sum of the weights of each dimension satisfies the constraints. The normalized weight vector is denoted as... .

[0036] After obtaining the initial fusion weights Next, the fusion weights need to be adjusted based on the risk type attributes of the insured individuals. Different risk types correspond to different information emphases. For example, for property risks, structured features (such as numerical attributes like building area, year of construction, and geographical location) often have higher risk discrimination value, while for health risks, the semantic information carried by unstructured features (such as medical records and image report descriptions) is more crucial. The risk type attributes are mapped to an adjustment factor vector through a pre-defined risk type coding table. Each element of the vector Indicates the first under the current risk type The weight scaling factor for each dimensional component. This will adjust the factor vector. With initial fusion weights Element-wise multiplication yields the constrained fusion weight vector. ,Right now To prevent extreme values ​​from appearing in the weights of certain dimensions after adjustment, [the following is done]: A truncation process is performed to limit weight values ​​that exceed preset upper and lower bounds to a reasonable range, thereby ensuring the numerical stability of the fusion process.

[0037] After completing the constraint adjustment of the fusion weights, the first mapping vector will be... With the second mapping vector according to By performing weighted combination, a unified representation vector is obtained. The weighted combination is calculated as follows: and Each dimension component is multiplied by its corresponding fusion weight, and then the two are added together. Specifically, let... The corresponding set of dimensional weight components is , The corresponding set of dimensional weight components is ,but The dimensional components are This weighted combination method, while preserving the semantic information of each modality, achieves fine-grained control over the contributions of different modalities through differentiated dimensional weights. This allows the unified representation vector to reflect both the precise attribute information of structured data and the rich semantic content contained in unstructured data.

[0038] The final unified representation vector As the core input for subsequent knowledge graph construction and insurance parameter reasoning, its dimension is consistent with the semantic space dimension. Because the allocation of fusion weights fully considers multi-granularity semantic correspondences and risk type constraints, It can adaptively emphasize different modalities of information in different types of insurance scenarios, thereby improving the accuracy of subsequent risk assessment and parameter configuration. In practical applications, when the risk type of the insured changes, only the adjustment factor vector needs to be updated. The value of can be used to quickly adjust the fusion strategy without retraining the projection matrix, which has strong flexibility and scalability.

[0039] Based on the unified representation vector and preset domain ontology rules, a knowledge graph containing entity nodes, attribute nodes, and relation edges of the insured object is determined. Starting from the entity node corresponding to the insured object, the knowledge graph is propagated through multi-hop relation paths to determine the set of associated entities and the set of associated attributes related to risk assessment, including: Based on the unified representation vector, the entity type of the object to be insured in the preset domain ontology rules is identified by semantic matching, and the entity node of the object to be insured is generated. Based on the feature dimension distribution in the unified representation vector, the attribute information of the object to be insured is extracted through semantic parsing, and the attribute information is labeled with type and hierarchically divided according to the attribute definition in the domain ontology rules to generate an attribute node set; Based on the relation type constraints defined in the domain ontology rules, relation edges are established between the entity nodes and the attribute node set to obtain the knowledge graph; Starting from the entity node corresponding to the insured object, a multi-hop traversal is performed along the relation edges. During the traversal, the propagation path is filtered according to the semantic type of the relation edges to obtain a set of candidate association paths. Extract entity nodes related to risk assessment semantics from the termination nodes of the candidate association path set to obtain the association entity set; The associated attribute set is obtained by extracting the attribute nodes related to risk quantification calculation from the attribute nodes associated with each entity node in the associated entity set.

[0040] When constructing a knowledge graph based on a unified representation vector and predefined domain ontology rules, the first step is to determine the entity type of the insured object within the domain ontology. The domain ontology rules predefine an entity type system in the insurance domain, such as natural persons, legal persons, vehicles, buildings, and medical entities. Each entity type corresponds to a set of semantic description vectors. By calculating the cosine similarity between the unified representation vector and the semantic description vectors of each entity type, the entity type with the highest similarity is selected as the entity type label for the insured object. Based on this, according to the node template specified for that entity type, entity nodes for the insured object are generated. Each node carries an entity type label, a unique identifier, and core semantic summary information derived from the unified representation vector, ensuring the traceability and semantic integrity of the entity nodes within the knowledge graph.

[0041] The generation of attribute nodes relies on the feature dimension distribution information of the unified representation vector. Different dimensional intervals of the unified representation vector carry semantic content from structured and unstructured data, respectively. For example, low-dimensional intervals correspond to structured attributes such as age, occupation, and region, while high-dimensional intervals correspond to unstructured attributes such as health status descriptions and historical behavioral characteristics extracted from text or images. Through semantic parsing, the unified representation vector is mapped according to its dimensional distribution to a predefined list of attribute definitions in the domain ontology rules, identifying the various attribute information possessed by the insured object. The identified attribute information is then labeled according to the attribute definitions in the domain ontology rules, such as classifying attributes into categories like basic identity attributes, risk behavior attributes, health status attributes, and asset status attributes. Furthermore, the attributes are hierarchically divided based on their semantic hierarchical relationships, forming a hierarchical set of attribute nodes. Each attribute node records the attribute name, attribute type label, attribute value, and its hierarchical information, providing a structured foundation for the subsequent establishment of relational edges.

[0042] The establishment of relation edges follows the predefined relation type constraints in the domain ontology rules. These rules define the allowed relation types between entity nodes and attribute nodes, such as "owns attribute," "belongs to," "influences," and "associates." Each relation type comes with semantic constraints, specifying the type requirements and attribute compatibility conditions for the nodes at both ends of the relation. Between the sets of entity nodes and attribute nodes, relation types that satisfy the constraints are matched based on the type labels and hierarchical information of each attribute node. Directed relation edges are established between the constrained entity node-attribute node pairs. Each relation edge carries a relation type label, a confidence score, and a source annotation. The confidence score reflects the reliability of the relation under the current unified representation vector semantics and is determined by the semantic matching score during attribute recognition. After these steps, a knowledge graph containing entity nodes, attribute nodes, and relation edges is constructed, forming a local semantic network structure centered on the insured object.

[0043] After the knowledge graph is constructed, starting with the entity node corresponding to the insured object, a multi-hop traversal operation is performed along the relation edges, gradually expanding outward to entities and attributes indirectly related to the insured object. During the multi-hop traversal, the propagation path of each hop is not extended indiscriminately, but is filtered according to the semantic type of the current relation edge. Specifically, a whitelist of relation types related to risk assessment semantics is pre-defined, such as "risk impact," "historical association," "similar association," and "environmental dependence," while relation types that purely describe administrative affiliation or are unrelated to business logic are excluded. During each hop propagation, the path only extends along the relation edges corresponding to the relation types in the whitelist, thus ensuring that the traversal path always maintains semantic consistency with the risk assessment target. The maximum number of hops for traversal is determined according to the path depth limit parameter in the domain ontology rules to prevent semantic drift and wasted computational resources due to excessively long paths. During the traversal, the complete path from the starting node to each reachable node is recorded, forming a candidate association path set. Each path in the set contains a sequence of nodes traversed and a sequence of relation edges.

[0044] Entity nodes semantically relevant to risk assessment are extracted from the terminal nodes of the candidate associated path set to form an associated entity set. The semantic relevance of a terminal node to risk assessment is determined by matching its entity type label against a pre-defined list of risk assessment entity types. For example, entity types such as medical institutions, accident records, occupational risk levels, and geographical risk areas are included in the risk assessment-related entity type list. If the same entity node is reached multiple times via multiple paths, the path record with the highest path confidence is retained, and the risk association strength of that entity node is cumulatively scored. Entity nodes with higher risk association strength have higher weight in subsequent reasoning. Each entity node in the associated entity set carries its corresponding path source information and risk association strength score, providing verifiable semantic support for the constrained reasoning stage.

[0045] The extraction of the associated attribute set is based on the associated entity set. For each entity node in the associated entity set, all associated attribute nodes in the knowledge graph are retrieved, and then attribute nodes relevant to risk quantification calculation are selected. The criteria for determining attribute nodes relevant to risk quantification calculation are based on matching the attribute node's type label with a pre-defined list of risk quantification attribute types. For example, attribute types such as accident frequency, severity of past medical history, property valuation, driving experience, and building fire resistance rating are all required attributes for risk quantification calculation. When multiple attribute nodes of the same attribute type exist under different entity nodes, the attribute nodes are weighted and sorted according to the risk association strength of the entity nodes, prioritizing the retention of attribute nodes under entities with high association strength to ensure the information quality of the associated attribute set and the accuracy of risk calculation. In the final associated attribute set, each attribute node records the attribute name, attribute value, attribute type, the identifier of the entity node to which it belongs, and the corresponding risk quantification weight, providing a complete input data structure for subsequently determining the insurance parameter configuration scheme through constrained reasoning.

[0046] The entire knowledge graph construction and multi-hop propagation process is carried out under the constraint framework of domain ontology rules. As a source of prior knowledge, the domain ontology rules ensure the semantic consistency of each stage of entity recognition, attribute extraction, relationship establishment and path selection, avoiding the semantic bias problem that may occur in small sample scenarios in pure data-driven methods. This ensures that the final set of associated entities and associated attributes can truly reflect the full picture of the risk characteristics of the insured object, laying a reliable graph knowledge foundation for the accurate generation of insurance parameter configuration schemes.

[0047] Based on the feature dimension distribution in the unified representation vector, the attribute information of the insured object is extracted through semantic parsing. The attribute information is then labeled with types and hierarchically divided according to the attribute definitions in the domain ontology rules, generating a set of attribute nodes, including: Semantic parsing is performed on the feature dimension distribution in the unified representation vector to identify the semantic concepts corresponding to each feature dimension, thereby obtaining a set of semantic concepts. The semantic concepts in the semantic concept set are semantically matched with the predefined attribute definitions in the domain ontology rules to determine the mapping relationship between semantic concepts and attribute types, and the corresponding attribute information is extracted for the insured object based on the mapping relationship. Based on the attribute inheritance and attribute dependency relationships defined in the domain ontology rules, hierarchical analysis is performed on the attribute information to determine the hierarchical relationships and dependency constraints between each attribute information. The attribute information is organized into a tree-like hierarchical structure according to the hierarchical relationship, and constraint markers are determined between the nodes of the tree-like hierarchical structure based on the dependency constraint relationship. The attribute information of each level in the tree hierarchy and the constraint markers are instantiated as attribute nodes to obtain the attribute node set.

[0048] After obtaining the unified representation vector, it is necessary to extract the attribute information of the insured object from it in order to construct the attribute nodes in the knowledge graph. The various feature dimensions of the unified representation vector are not randomly distributed, but carry semantic information from the fusion of structured and unstructured data. Semantic parsing of these feature dimensions is essentially mapping the numerical distribution in the vector space back to interpretable semantic concepts, thereby providing a foundation for the integration of domain ontology rules.

[0049] When performing semantic parsing on the feature dimension distribution in the unified representation vector, the vector is divided into several sub-intervals according to a preset semantic partitioning scheme. Each sub-interval corresponds to a predefined semantic slot. The division of semantic slots is based on statistical analysis of a large number of insurance samples during the training phase. There is a stable correspondence between different dimension intervals and specific semantic concepts. For example, the activation intensity of one interval mainly reflects health status-related information, while another interval mainly reflects asset size-related information. By calculating the activation intensity distribution of each dimension interval, significantly activated semantic slots are identified, thus obtaining a set of semantic concepts. Each semantic concept in the set of semantic concepts is accompanied by a confidence score. Semantic concepts with confidence scores below a preset threshold are filtered out to ensure the accuracy of subsequent processing.

[0050] When semantically matching each semantic concept in the semantic concept set with the predefined attribute definitions in the domain ontology rules, a similarity calculation method based on semantic embedding is adopted. Each attribute definition in the domain ontology rules is pre-encoded as an attribute embedding vector, and each semantic concept in the semantic concept set is also encoded as a concept embedding vector. By calculating the similarity between the concept embedding vector and each attribute embedding vector, the optimal mapping relationship from semantic concept to attribute type is determined. When the similarity between a semantic concept and multiple attribute definitions exceeds the matching threshold, disambiguation is performed according to the attribute priority order in the domain ontology rules, prioritizing the attribute type with a higher matching degree to the business scenario of the insured object. After the mapping relationship is determined, the values ​​of the corresponding dimension range are extracted from the unified representation vector according to the mapping relationship. Combined with the attribute type value specification, the values ​​are converted into specific values ​​of attribute information, completing the extraction of attribute information.

[0051] The hierarchical analysis of attribute information relies on the attribute inheritance and dependency relationships explicitly defined in the domain ontology rules. Attribute inheritance describes the hierarchical semantic inclusion relationship between attributes; for example, "chronic medical history" is a subordinate attribute of "health status," and "health status" is a subordinate attribute of "personal risk characteristics." Attribute dependency describes the logical constraints between attributes; for example, the validity of "past claims records" depends on the existence of the "insurance contract validity period" attribute. The extracted attribute information set is iterated through the attribute inheritance definition in the domain ontology rules, determining the position of the attribute type corresponding to each attribute information in the inheritance graph, and identifying the superior attribute link for each attribute, thereby establishing the hierarchical relationship between attribute information. Simultaneously, based on the attribute dependency definition, it checks whether the attribute pairs in the current attribute information set satisfy dependency constraints, recording the attribute pairs with dependency relationships and their constraint types. Constraint types include necessary dependency, mutual exclusion constraint, and conditional dependency. Necessary dependency means that the existence of one attribute is predicated on the existence of another attribute; mutual exclusion constraint means that two attributes cannot both take non-null values; and conditional dependency means that the value range of one attribute is restricted by the value range of another attribute.

[0052] When organizing attribute information into a tree-like hierarchical structure according to hierarchical relationships, the top-level attribute category in the domain ontology rules is used as the root node, expanding downwards layer by layer. The depth of the tree-like hierarchical structure is determined by the maximum number of levels in the attribute inheritance relationship in the domain ontology rules, typically set to 3 to 5 levels in the insurance domain. When an attribute has multiple superior paths in the inheritance relationship, the superior path with the shortest path length is selected as the main path, and the attribute is attached to the nearest superior attribute node to avoid duplicate attachments in the tree structure. After the tree-like hierarchical structure is constructed, dependency constraint markers are superimposed between the nodes. For necessary dependencies, a necessary dependency constraint marker is added between the dependent node and the dependent node; for mutually exclusive constraints, a mutually exclusive constraint marker is added between two mutually exclusive nodes; for conditional dependencies, a conditional dependency constraint marker is added between the relevant node pairs, along with the specific content of the constraint condition, such as "when the value of the 'Previous Claim Amount' node is greater than a certain threshold, the value range of the 'Risk Level' node is narrowed to the high-risk range." Constraint tags are attached to the edges of the tree hierarchy as structured annotations. They do not change the topological relationship of the tree structure, but can be directly read by the constraint inference engine in subsequent inference stages.

[0053] When instantiating attribute information and constraint tags at each level of the tree hierarchy into attribute nodes, the data structure of each attribute node includes five fields: attribute identifier, attribute type label, attribute value, level number, parent node reference, and constraint tag list. The attribute identifier is generated by concatenating the attribute type code with the unique identifier of the insured object, ensuring the attribute node's global uniqueness throughout the knowledge graph. The attribute type label is directly derived from the attribute type definition in the domain ontology rules, ensuring that the attribute node's type semantics are consistent with the ontology rules. Before being written to the attribute node, the attribute value undergoes validity verification based on the attribute type's value range constraints. Attribute values ​​that do not meet the value range constraints are marked as abnormal and trigger a manual review process to prevent erroneous data from polluting the knowledge graph. The level number records the depth position of the attribute node in the tree hierarchy, facilitating filtering of attribute nodes by level during subsequent multi-hop relationship path propagation. The parent node reference records the identifier of the directly parent attribute node, supporting backtracking from any attribute node along the inheritance relationship to the root node. The constraint tag list summarizes all the dependency constraints that the node participates in. Each constraint tag includes the constraint type, the associated node identifier, and a description of the constraint condition. After all attribute nodes are instantiated, they are aggregated to form an attribute node set. This set, along with the entity node set, serves as input for knowledge graph construction, supporting the subsequent determination of the associated entity set and associated attribute set.

[0054] Based on the set of associated entities, the set of associated attributes, and the unified representation vector, an insurance parameter configuration scheme is determined through constraint reasoning. Execution feedback data for the insurance parameter configuration scheme is obtained, and the parameters in the cross-modal alignment process are adaptively adjusted using this execution feedback data, including: Knowledge graph logic rules are extracted from the set of associated entities and the set of associated attributes, and the semantic similarity between each associated entity in the set of associated entities and the insured object is calculated based on the unified representation vector to obtain the semantic similarity distribution; Based on the logical rules of the knowledge graph, the value space of candidate insurance parameters is constrained and filtered to obtain a set of compliant parameter values. Then, the combination of parameter values ​​in the set of compliant parameter values ​​is prioritized according to the semantic similarity distribution to obtain the insurance parameter configuration scheme. Obtain execution feedback data of the insurance parameter configuration scheme in actual insurance business, and calculate the deviation between the insurance parameter configuration scheme and actual business needs based on the execution feedback data; The feature vector causing the deviation in the multimodal data is determined based on the deviation amount, and the parameters in the cross-modal alignment process are adaptively adjusted based on the feature vector causing the deviation.

[0055] When extracting logical rules from the set of associated entities and the set of associated attributes, it is necessary to traverse all associated entities and their attribute nodes related to the insured object in the knowledge graph, and identify the rule structures with reasoning constraints. These logical rules are usually in the form of condition-conclusion, such as "If the industry category of the associated entity belongs to a high-risk industry and the historical loss ratio exceeds the threshold, then the lower limit of the corresponding premium parameter should be raised to a certain range." During the extraction of logical rules, it is also necessary to verify the validity of the rules in conjunction with the domain ontology rules, filtering out rules that are irrelevant to the current insurance scenario or whose confidence level is lower than the preset threshold, and retaining the set of rules that have substantial constraints on the configuration of insurance parameters.

[0056] When calculating semantic similarity, a unified representation vector is used as the semantic anchor of the insured object. A corresponding semantic vector representation is also constructed for each associated entity in the associated entity set. The semantic vector of an associated entity can be obtained by weighted aggregation of its attribute node embeddings and relation edge weights. Let the unified representation vector of the insured object be... , No. The semantic vectors of the associated entities are Then the semantic similarity between the two Calculated using cosine similarity: ; After calculating all related entities, each Normalization generates a semantic similarity distribution, which is used to prioritize the combination of subsequent parameter values. The semantic similarity distribution can reflect the reference value of different related entities for the risk characteristics of the insured object. The higher the similarity of the related entity, the stronger the reference significance of its historical insurance parameters for the current configuration scheme.

[0057] Based on the extracted knowledge graph logical rules, the value space of candidate insurance parameters is constrained and filtered. This value space is predefined by the business system and includes a set of discrete or continuous values ​​across multiple dimensions, such as premium range, deductible range, and coverage period. The constraint filtering process applies logical rules one by one, removing parameter value combinations that do not satisfy any mandatory constraint rule from the candidate space, ultimately resulting in a set of compliant parameter values. For rules with soft constraints (i.e., advisory rather than mandatory constraints), the corresponding values ​​are not directly eliminated; instead, their priority weight is reduced in subsequent ranking stages.

[0058] After the set of compliance parameter values ​​is determined, the combinations of parameter values ​​are prioritized according to the semantic similarity distribution. Let the first one be the first one. The comprehensive score for the combination of parameter values ​​is It consists of two parts: one is a weighted reference score based on the semantic similarity distribution, that is, based on the semantic similarity of each associated entity. The system uses two weighted averages: first, a weighted summation of the matching degree between the corresponding parameter values ​​of the associated entity in historical cases and the current candidate combinations; and second, a soft constraint satisfaction score, reflecting the degree to which the parameter value combination conforms to the suggested rules. These two scores are then combined according to a preset ratio to obtain the final score for each parameter value combination. The combinations are then sorted from highest to lowest score, and the top-ranked combinations are selected to form the insurance parameter configuration scheme. The optimal combination is then output as the recommended scheme.

[0059] During the actual insurance underwriting process, execution feedback data of the insurance parameter configuration scheme is collected. This feedback data includes, but is not limited to: actual underwriting results (successful or rejected coverage), customer acceptance of the parameter configuration, subsequent claims incidence and payout amounts, and manual review comments from sales personnel on the configuration scheme. This feedback data reflects the degree of fit between the configuration scheme and actual business needs from different dimensions and requires structured processing before it can be used for deviation calculation.

[0060] Based on execution feedback data, the deviation between the insurance parameter configuration scheme and actual business needs is calculated. Let the deviation be the first parameter in the configuration scheme. The recommended values ​​for each parameter are: The actual reasonable value for this parameter, determined based on the execution feedback data, is... Then the deviation of this parameter Defined as: ; After calculating the deviation for all parameters, the deviations are summarized to form a deviation vector. The absolute value of each component reflects the degree of deviation of the corresponding parameter. When the deviation of a parameter exceeds a preset tolerance threshold, the subsequent feature vector tracing and parameter adjustment process is triggered. The calculation of deviation is not only for numerical parameters, but also for categorical parameters, the degree of deviation can be quantified by category matching degree or business rule compliance rate.

[0061] Based on the deviation vector To identify the feature vectors causing bias in multimodal data, a correlation analysis is performed between the bias vectors and the contribution records of each feature dimension during the cross-modal alignment process. In the cross-modal alignment stage, a traceable computational relationship exists between each output dimension and both the input structured and unstructured feature vectors. Using gradient backpropagation or feature attribution methods, the sensitivity of the bias to each input feature is calculated; feature dimensions with higher sensitivity are identified as the feature vectors causing the bias. Let the first... Each input feature pairs the bias The attributable contribution value Then the comprehensive attribution score of this feature over all parameter deviations. The overall attribution score is obtained by weighted summation of all parameters, with the weights proportional to the absolute value of the deviation of the corresponding parameter. The higher the value, the greater the impact of this feature on the overall deviation, and it needs to be given special attention when adjusting the alignment parameters.

[0062] After identifying the feature vectors causing the bias, the relevant parameters in the cross-modal alignment process are adaptively adjusted. The adjustable parameters mainly include: the projection matrix parameters of structured and unstructured features, the association strength values ​​in the multi-granularity correspondence matrix, and the weight allocation of each dimension in the fusion weight vector. For feature dimensions with high attribution scores, their proportion in the fusion weight is appropriately reduced, or the parameter values ​​of the corresponding rows and columns in their projection matrices are adjusted to reduce the excessive influence of these features on subsequent configuration results. The adjustment magnitude is proportional to the amount of bias, and an upper limit is set for the adjustment step size to avoid drastic fluctuations in the alignment results due to excessively large single adjustments.

[0063] After adaptive adjustment, the multimodal data of the current insured object is reprocessed using the adjusted cross-modal alignment parameters to generate a new unified representation vector. The knowledge graph reasoning and constraint filtering process is then re-executed to verify whether the adjusted insurance parameter configuration scheme has a lower bias on the known feedback samples. If the verification passes, the adjusted parameters are persistently saved as initial parameters for subsequent insurance parameter configuration tasks. If the verification fails, the system reverts to the parameter state before adjustment, and the reason for the failure is recorded for subsequent optimization analysis. Through multiple rounds of feedback and parameter adjustment iterations, the parameters of the cross-modal alignment process gradually converge to a state that better matches actual business needs, thereby continuously improving the accuracy and compliance of the insurance parameter configuration scheme.

[0064] Based on the logical rules of the knowledge graph, the value space of candidate insurance parameters is constrained and filtered to obtain a set of compliant parameter values. Then, according to the semantic similarity distribution, the combinations of parameter values ​​in the compliant parameter value set are prioritized to obtain the insurance parameter configuration scheme, including: The logical rules of the knowledge graph are parsed to obtain a set of parameter boundary conditions and a set of parameter consistency conditions. The set of parameter boundary conditions limits the upper and lower bounds of the value of a single parameter, and the set of parameter consistency conditions limits the value association relationship between multiple parameters. The compliance of each parameter value combination in the candidate insurance parameter value space is verified. It is determined whether the parameter value combination simultaneously satisfies the parameter boundary condition set and the parameter consistency condition set. The parameter value combinations that satisfy all conditions are retained and the parameter value combinations that violate any condition are removed to obtain the compliant parameter value set. For each combination of parameter values ​​in the set of compliance parameter values, a comprehensive semantic similarity score is calculated between the associated entity corresponding to the parameter value combination and the insured object based on the semantic similarity distribution. The parameter value combinations in the compliance parameter value set are sorted in descending order according to the comprehensive semantic similarity score, and the parameter value combination with the highest sorted position is selected as the insurance parameter configuration scheme.

[0065] After completing multi-hop relationship path propagation and obtaining the sets of associated entities and attributes, the domain logic rules accumulated in the knowledge graph need to be transformed into operable constraints to systematically screen the value space of candidate insurance parameters. Knowledge graph logic rules are typically stored in the rule layer of the graph in the form of predicate logic or conditional expressions. The parsing process needs to split these rules into two independent sets of conditions according to constraint type: parameter boundary condition set and parameter consistency condition set. Each rule in the parameter boundary condition set describes the legal value range of a single insurance parameter, such as the lower limit of the insured amount for a certain type of property insurance not being lower than the appraised value of the insured property multiplied by a specific coefficient, or the upper limit of the insurance period being limited by the insured's age. Each rule in the parameter consistency condition set describes the coupling constraint relationship between multiple parameters, such as the ratio between the deductible and the insured amount not exceeding the regulatory range, or the ratio between the supplementary insurance premium and the main insurance premium meeting actuarial consistency requirements. During parsing, each logical expression in the rule layer is subjected to syntactic analysis to identify the parameter identifiers, comparison operators and logical connectors involved. Single-parameter constraints are classified into the parameter boundary condition set, and cross-parameter constraints are classified into the parameter consistency condition set. These are then stored in a structured form for easy verification and retrieval later.

[0066] The compliance verification process is performed on each combination of parameter values ​​in the candidate insurance parameter value space. The verification process is executed sequentially in two stages: The first stage targets the parameter boundary condition set, checking whether the specific value of each parameter in the value combination falls within the legal upper and lower bounds of the corresponding parameter. If a parameter value violates its boundary condition, the value combination is directly marked as non-compliant and skipped from subsequent checks. The second stage targets the parameter consistency condition set. For value combinations that pass the first stage verification, the parameter values ​​involving cross-parameter constraints are extracted and substituted into the corresponding consistency condition expressions for logical judgment. If any consistency condition is not met, the value combination is also marked as non-compliant. Only when a parameter value combination passes both the boundary condition verification and the consistency condition verification is it retained in the compliant parameter value set. This two-stage sequential verification strategy can quickly eliminate a large number of obviously non-compliant combinations in the first stage using boundary conditions, thereby reducing the computational overhead of the consistency condition verification in the second stage and improving overall screening efficiency.

[0067] After obtaining the set of compliance parameter values, a comprehensive semantic similarity score needs to be calculated for each combination of parameter values ​​to quantify the semantic matching degree between the combination and the insured object. Each combination of parameter values ​​corresponds to several associated entities in the knowledge graph. These associated entities are connected to the entity nodes of the insured object through multi-hop paths, and the semantic similarity distribution between each associated entity and the insured object has been calculated in the previous steps. For a given combination of parameter values, all associated entities corresponding to it are collected. Let the number of associated entities corresponding to this combination be . , No. The semantic similarity between each associated entity and the insured object is The path association weight between this associated entity and the current parameter value combination in the knowledge graph is: The combined semantic similarity score of this parameter's value combination is then calculated. Calculated using weighted aggregation: ; Among them, path association weight The weighted aggregation method determines the path association weight based on the hop distance and edge type importance between associated entities and parameter value combinations. The shorter the hop distance and the higher the correlation between the edge type and the risk assessment, the greater the corresponding path association weight. This weighted aggregation method can fully utilize the differentiated information contained in the semantic similarity distribution and avoid diluting the contribution of highly relevant entities due to simple averaging.

[0068] After calculating the comprehensive semantic similarity score for all parameter value combinations in the compliance parameter value set, they are sorted in descending order of score. The sorting results reflect the order of semantic matching degree between each parameter value combination and the insured object from high to low. The higher-ranked combination means that its associated entity group is closer to the feature description of the insured object, and has stronger supporting evidence in terms of historical cases and domain knowledge. The parameter value combination with the highest ranking position is selected as the final output insurance parameter configuration scheme. This scheme not only satisfies all compliance constraints stipulated by the knowledge graph logic rules, but also achieves optimal matching with the insured object in the semantic similarity dimension, thus achieving a balance between compliance and personalized adaptability.

[0069] In practical applications, when the dimensionality of the candidate insurance parameters' value space is high, the number of parameter value combinations grows exponentially, making the computational cost of directly enumerating all combinations for verification and sorting substantial. To address this, a pruning strategy can be introduced during the compliance verification phase: in the first stage of screening based on parameter boundary conditions, a parameter-by-parameter approach is used to construct value combinations. Parameter values ​​that violate boundary conditions are no longer combined with other parameter values ​​for expansion, thus pruning a large number of invalid combinations during the combination construction phase. In the consistency condition verification phase, the most stringent constraints are prioritized for verification, identifying and eliminating non-compliant combinations as quickly as possible. Furthermore, the calculation of the comprehensive semantic similarity score can be performed in parallel with compliance verification. Combinations that pass verification are promptly scored and inserted into an ordered queue, avoiding the additional time overhead of unified sorting after all verifications are completed. Through these engineering optimizations, the entire constraint reasoning and priority sorting process can be completed within an acceptable response time, meeting the real-time requirements of actual insurance business.

[0070] The final output of the insurance parameter configuration scheme is presented in a structured form, including the specific recommended value for each insurance parameter and its corresponding comprehensive semantic similarity score. It also includes a list of key related entities and path information supporting the recommendation, which facilitates business personnel to review and interpret the configuration results.

[0071] A second aspect of this invention provides an insurance parameter configuration system based on multimodal data and knowledge graphs, comprising: The feature extraction unit is used to acquire multimodal data of the insured object, extract features from the multimodal data, and obtain structured feature vectors and unstructured feature vectors. The alignment processing unit is used to perform cross-modal alignment processing on the structured feature vector and the unstructured feature vector to obtain a unified representation vector. The graph construction unit is used to determine a knowledge graph containing entity nodes, attribute nodes and relation edges of the insured object based on the unified representation vector and the preset domain ontology rules. In the knowledge graph, starting from the entity node corresponding to the insured object, the set of associated entities and the set of associated attributes related to risk assessment are determined through multi-hop relation path propagation. The configuration adjustment unit is used to determine the insurance parameter configuration scheme through constraint reasoning based on the associated entity set, the associated attribute set, and the unified representation vector, obtain the execution feedback data of the insurance parameter configuration scheme, and use the execution feedback data to adaptively adjust the parameters in the cross-modal alignment process.

[0072] A third aspect of the present invention provides an electronic device, comprising: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0073] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0074] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.

[0075] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for configuring insurance parameters based on multimodal data and knowledge graphs, characterized in that, include: Obtain multimodal data of the insured object, and extract features from the multimodal data to obtain structured feature vectors and unstructured feature vectors; The structured feature vector and the unstructured feature vector are aligned across modes to obtain a unified representation vector. Based on the unified representation vector and the preset domain ontology rules, a knowledge graph containing entity nodes, attribute nodes, and relation edges of the insured object is determined. Starting from the entity node corresponding to the insured object in the knowledge graph, the set of associated entities and the set of associated attributes related to risk assessment are determined through multi-hop relation path propagation. Based on the set of associated entities, the set of associated attributes, and the unified representation vector, an insurance parameter configuration scheme is determined through constraint reasoning. Execution feedback data of the insurance parameter configuration scheme is obtained, and the parameters in the cross-modal alignment process are adaptively adjusted using the execution feedback data.

2. The method according to claim 1, characterized in that, Obtain multimodal data of the insured object, and perform feature extraction on the multimodal data to obtain structured feature vectors and unstructured feature vectors, including: The structured attribute data and unstructured descriptive data of the insured object are obtained as the multimodal data, wherein the structured attribute data is represented by key-value pairs and the unstructured descriptive data is represented by text sequences; The structured attribute data is subjected to field standardization and missing value imputation to obtain standardized structured data; the unstructured description data is subjected to word segmentation and stop word filtering to obtain standardized unstructured data. The numerical fields in the standardized structured data are mapped to a unified numerical space through normalization transformation, and the categorical fields in the standardized structured data are converted into vector representations through encoding mapping. The normalized result of the numerical field is concatenated with the vector representation of the categorical field to obtain the structured feature vector; The normalized unstructured data is input into a pre-trained semantic representation network to extract semantic feature representations from the normalized unstructured data. The semantic feature representations are then pooled and aggregated to obtain the unstructured feature vector.

3. The method according to claim 1, characterized in that, The structured feature vector and the unstructured feature vector are subjected to cross-modal alignment to obtain a unified representation vector, including: Semantic space projection processing is performed on the structured feature vector and the unstructured feature vector respectively, mapping the structured feature vector to obtain a first mapping vector and the unstructured feature vector to obtain a second mapping vector; Calculate the similarity distribution between the first mapping vector and the second mapping vector at multiple semantic granularity levels to determine a multi-granularity cross-modal correspondence matrix. The multi-granularity cross-modal correspondence matrix simultaneously characterizes the correlation strength between the structured feature vector and the unstructured feature vector at both the global and local semantic levels. Based on the multi-granularity cross-modal correspondence matrix, fusion weights are assigned to each dimension component of the first mapping vector and the second mapping vector, and the fusion weights are constrained and adjusted according to the risk type attribute of the insured object; The first mapping vector and the second mapping vector are weighted and combined according to the fusion weights adjusted by the constraints to obtain the unified representation vector.

4. The method according to claim 1, characterized in that, Based on the unified representation vector and preset domain ontology rules, a knowledge graph containing entity nodes, attribute nodes, and relation edges of the insured object is determined. Starting from the entity node corresponding to the insured object, the knowledge graph is propagated through multi-hop relation paths to determine the set of associated entities and the set of associated attributes related to risk assessment, including: Based on the unified representation vector, the entity type of the object to be insured in the preset domain ontology rules is identified by semantic matching, and the entity node of the object to be insured is generated. Based on the feature dimension distribution in the unified representation vector, the attribute information of the object to be insured is extracted through semantic parsing, and the attribute information is labeled with type and hierarchically divided according to the attribute definition in the domain ontology rules to generate an attribute node set; Based on the relation type constraints defined in the domain ontology rules, relation edges are established between the entity nodes and the attribute node set to obtain the knowledge graph; Starting from the entity node corresponding to the insured object, a multi-hop traversal is performed along the relation edges. During the traversal, the propagation path is filtered according to the semantic type of the relation edges to obtain a set of candidate association paths. Extract entity nodes related to risk assessment semantics from the termination nodes of the candidate association path set to obtain the association entity set; The associated attribute set is obtained by extracting the attribute nodes related to risk quantification calculation from the attribute nodes associated with each entity node in the associated entity set.

5. The method according to claim 4, characterized in that, Based on the feature dimension distribution in the unified representation vector, the attribute information of the insured object is extracted through semantic parsing. The attribute information is then labeled with types and hierarchically divided according to the attribute definitions in the domain ontology rules, generating a set of attribute nodes, including: Semantic parsing is performed on the feature dimension distribution in the unified representation vector to identify the semantic concepts corresponding to each feature dimension, thereby obtaining a set of semantic concepts. The semantic concepts in the semantic concept set are semantically matched with the predefined attribute definitions in the domain ontology rules to determine the mapping relationship between semantic concepts and attribute types, and the corresponding attribute information is extracted for the insured object based on the mapping relationship. Based on the attribute inheritance and attribute dependency relationships defined in the domain ontology rules, hierarchical analysis is performed on the attribute information to determine the hierarchical relationships and dependency constraints between each attribute information. The attribute information is organized into a tree-like hierarchical structure according to the hierarchical relationship, and constraint markers are determined between the nodes of the tree-like hierarchical structure based on the dependency constraint relationship. The attribute information of each level in the tree hierarchy and the constraint markers are instantiated as attribute nodes to obtain the attribute node set.

6. The method according to claim 1, characterized in that, Based on the set of associated entities, the set of associated attributes, and the unified representation vector, an insurance parameter configuration scheme is determined through constraint reasoning. Execution feedback data for the insurance parameter configuration scheme is obtained, and the parameters in the cross-modal alignment process are adaptively adjusted using this execution feedback data, including: Knowledge graph logic rules are extracted from the set of associated entities and the set of associated attributes, and the semantic similarity between each associated entity in the set of associated entities and the insured object is calculated based on the unified representation vector to obtain the semantic similarity distribution; Based on the logical rules of the knowledge graph, the value space of candidate insurance parameters is constrained and filtered to obtain a set of compliant parameter values. Then, the combination of parameter values ​​in the set of compliant parameter values ​​is prioritized according to the semantic similarity distribution to obtain the insurance parameter configuration scheme. Obtain execution feedback data of the insurance parameter configuration scheme in actual insurance business, and calculate the deviation between the insurance parameter configuration scheme and actual business needs based on the execution feedback data; The feature vector causing the deviation in the multimodal data is determined based on the deviation amount, and the parameters in the cross-modal alignment process are adaptively adjusted based on the feature vector causing the deviation.

7. The method according to claim 6, characterized in that, Based on the logical rules of the knowledge graph, the value space of candidate insurance parameters is constrained and filtered to obtain a set of compliant parameter values. Then, according to the semantic similarity distribution, the combinations of parameter values ​​in the compliant parameter value set are prioritized to obtain the insurance parameter configuration scheme, including: The logical rules of the knowledge graph are parsed to obtain a set of parameter boundary conditions and a set of parameter consistency conditions. The set of parameter boundary conditions limits the upper and lower bounds of the value of a single parameter, and the set of parameter consistency conditions limits the value association relationship between multiple parameters. The compliance of each parameter value combination in the candidate insurance parameter value space is verified. It is determined whether the parameter value combination simultaneously satisfies the parameter boundary condition set and the parameter consistency condition set. The parameter value combinations that satisfy all conditions are retained and the parameter value combinations that violate any condition are removed to obtain the compliant parameter value set. For each combination of parameter values ​​in the set of compliance parameter values, a comprehensive semantic similarity score is calculated between the associated entity corresponding to the parameter value combination and the insured object based on the semantic similarity distribution. The parameter value combinations in the compliance parameter value set are sorted in descending order according to the comprehensive semantic similarity score, and the parameter value combination with the highest sorted position is selected as the insurance parameter configuration scheme.

8. An insurance parameter configuration system based on multimodal data and knowledge graphs, used to implement the method as described in any one of claims 1-7, characterized in that, include: The feature extraction unit is used to acquire multimodal data of the insured object, extract features from the multimodal data, and obtain structured feature vectors and unstructured feature vectors. The alignment processing unit is used to perform cross-modal alignment processing on the structured feature vector and the unstructured feature vector to obtain a unified representation vector. The graph construction unit is used to determine a knowledge graph containing entity nodes, attribute nodes and relation edges of the insured object based on the unified representation vector and the preset domain ontology rules. In the knowledge graph, starting from the entity node corresponding to the insured object, the set of associated entities and the set of associated attributes related to risk assessment are determined through multi-hop relation path propagation. The configuration adjustment unit is used to determine the insurance parameter configuration scheme through constraint reasoning based on the associated entity set, the associated attribute set, and the unified representation vector, obtain the execution feedback data of the insurance parameter configuration scheme, and use the execution feedback data to adaptively adjust the parameters in the cross-modal alignment process.

9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 7.