A data representation method, device, apparatus and storage medium

CN122597749APending Publication Date: 2026-08-18PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610647706.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-11
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

例如,车辆车身整洁度的分析,主要依赖于人工检查或基于单一模态的图像识别技术,前者严重依赖人工主观性强、也存在分析效率低下的缺点,后者存在所依赖的分析数据模态单一、不够全面性的缺陷

Benefits of technology

本申请所述的数据表征方法,通过获取目标数据端上传的信息提取材料;提取所述信息提取材料中的文本特征信息和图片特征信息;融合所述文本特征信息和所述图片特征信息,获得融合特征信息;对所述融合特征信息进行目标视觉特征和目标语义特征抽取,获得特征抽取结果;根据所述特征抽取结果和预设的特征区分策略,筛选出期望特征结果;基于所述期望特征结果和预设的级别划分策略,确定所述信息提取材料的表征等级,并对目标对象进行等级化表征,其中,所述信息提取材料为所述目标对象的信息提取材料。在金融科技技术领域车险理赔场景下,采用该方法能够对获取到的车辆事故图片和事故图片的描述文本材料进行特征提取、融合和期望特征结果筛选,采用多模态数据分析方式,进行车辆理赔数据分析,并对目标车辆的数据分析结果进行等级化表征,丰富数据分析方式,提高目标数据的分析准确率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122597749A_ABST
    Figure CN122597749A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of data analysis, and relates to a data representation method, device, equipment and storage medium. Information extraction materials uploaded by a target data terminal are acquired; text feature information and picture feature information are extracted; fusion feature information is obtained; target visual feature and target semantic feature extraction is performed on the fusion feature information; expected feature results are screened out; based on the expected feature results, the representation level of the information extraction materials is determined, and the target object is represented in a graded manner. In the vehicle insurance claim scene in the field of financial technology, the method can extract features, fuse and screen expected feature results from the acquired vehicle accident pictures and the description text materials of the accident pictures, adopt a multi-modal data analysis method, analyze vehicle claim data, and represent the data analysis results of the target vehicle in a graded manner, thereby enriching the data analysis method and improving the analysis accuracy of the target data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data analysis technology and is applied to the multimodal data review and characterization scenario in auto insurance claims. It relates to a data characterization method, device, equipment, and storage medium. Background Technology

[0002] With the continuous increase in car ownership, there are growing challenges in analyzing and identifying automotive-related business data. Traditional data analysis mainly relies on factors such as driver driving records, vehicle type, and usage, but these methods often neglect data analysis of the vehicle's actual condition. For example, the analysis of vehicle body cleanliness mainly relies on manual inspection or image recognition technology based on a single modality. The former is heavily dependent on human subjectivity and suffers from low analysis efficiency, while the latter relies on a single modality of analytical data and lacks comprehensiveness. Therefore, currently, there is a lack of a more intelligent and accurate method for analyzing vehicle body condition. Summary of the Invention

[0003] The purpose of this application is to provide a data characterization method, apparatus, device, and storage medium for more intelligent and accurate vehicle body condition analysis.

[0004] In a first aspect, embodiments of this application provide a data characterization method, which employs the following technical solution: A data representation method includes the following steps: The information extraction materials uploaded by the target data terminal are obtained, wherein the information extraction materials include text information extraction materials and image information extraction materials; Extract textual and image features from the information material; By fusing the text feature information and the image feature information, fused feature information is obtained; The fused feature information is used to extract target visual features and target semantic features to obtain feature extraction results; Based on the feature extraction results and the preset feature discrimination strategy, the desired feature results are selected; Based on the expected feature results and the preset level division strategy, the representation level of the information extraction material is determined, and the target object is represented by a level. The information extraction material is the information extraction material of the target object.

[0005] Secondly, embodiments of this application also provide a data characterization device, which adopts the technical solution described below: A data characterization device, comprising: The information extraction material acquisition module is used to acquire the information extraction material uploaded by the target data terminal, wherein the information extraction material includes text information extraction material and image information extraction material; The feature information extraction module is used to extract text feature information and image feature information from the information extraction material; The feature information fusion module is used to fuse the text feature information and the image feature information to obtain fused feature information; The target feature extraction module is used to extract target visual features and target semantic features from the fused feature information to obtain feature extraction results. The desired feature filtering module is used to filter out desired feature results based on the feature extraction results and a preset feature differentiation strategy. The hierarchical representation module is used to determine the representation level of the information extraction material based on the expected feature results and the preset level division strategy, and to perform hierarchical representation of the target object, wherein the information extraction material is the information extraction material of the target object.

[0006] Thirdly, embodiments of this application also provide a computer device that adopts the technical solution described below: A computer device includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the data representation method described above.

[0007] Fourthly, embodiments of this application also provide a computer-readable storage medium, which adopts the technical solutions described below: A computer-readable storage medium storing computer-readable instructions that, when executed by a processor, implement the steps of the data representation method described above.

[0008] Compared with the prior art, the embodiments of this application have the following main advantages: The data representation method described in this application involves: acquiring information extraction materials uploaded by the target data terminal; extracting textual and image feature information from the information extraction materials; fusing the textual and image feature information to obtain fused feature information; extracting target visual and semantic features from the fused feature information to obtain feature extraction results; filtering desired feature results based on the feature extraction results and a preset feature differentiation strategy; determining the representation level of the information extraction materials based on the desired feature results and a preset level division strategy, and performing hierarchical representation of the target object, wherein the information extraction materials are the information extraction materials of the target object. In the context of auto insurance claims in the field of financial technology, this method can perform feature extraction, fusion, and desired feature result filtering on acquired vehicle accident images and descriptive text materials. It employs multimodal data analysis to analyze vehicle claims data and performs hierarchical representation of the data analysis results for the target vehicle, enriching the data analysis methods and improving the accuracy of target data analysis. Attached Figure Description

[0009] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 This is an exemplary system architecture diagram to which this application can be applied; Figure 2 This is a flowchart of an embodiment of a data characterization method according to this application; Figure 3 The flowchart shown is a specific embodiment of the data representation method described in this application for pre-training a large multimodal recognition model; Figure 4 yes Figure 2 A flowchart of a specific embodiment of step 204 shown; Figure 5 yes Figure 2 A flowchart of a specific embodiment of step 205 shown; Figure 6 yes Figure 2 A flowchart of a specific embodiment of step 206 shown; Figure 7 This is a schematic diagram of one embodiment of a data characterization device according to this application; Figure 8 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation

[0011] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.

[0012] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0013] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0014] like Figure 1 As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables.

[0015] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.

[0016] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, and a desktop computer, etc.

[0017] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.

[0018] It should be noted that the data representation method provided in this application embodiment is generally executed by a server, and correspondingly, a data representation device is generally set in the server.

[0019] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0020] Continue to refer to Figure 2 The diagram illustrates a flowchart of an embodiment of a data characterization method according to this application. The data characterization method includes the following steps: Step 201: Obtain the information extraction materials uploaded by the target data terminal, wherein the information extraction materials include text information extraction materials and image information extraction materials.

[0021] Specifically, the target data end is, for example, a multimodal data upload end; the information extraction material includes multimodal information extraction material, more specifically, in the vehicle claims review scenario, the upload end for vehicle images and image description text after a vehicle accident; at this time, the text information extraction material includes the image description text, and the image information extraction material includes the vehicle image; assuming the vehicle image is an image of the impact point on the left side of the vehicle, and the image description text is: there is a dent of 5cm by 5cm in area below the left rear door of the vehicle, and there is paint peeling around the dent, and there are certain scratches and stains in front of the dent.

[0022] In this embodiment, upon receiving a target material extraction instruction, the information extraction material uploaded by the target data terminal is obtained through a preset data receiving interface. This acquisition of the information extraction material uploaded by the target data terminal can be achieved through the data receiving interface of a pre-trained multimodal recognition model. The information extraction material uploaded by the target data terminal is directly sent to the pre-trained multimodal recognition model, which then processes the information extraction material.

[0023] In this embodiment, by acquiring the information extraction materials uploaded by the target data terminal, the multimodal data extraction materials uploaded by the multimodal data upload terminal are obtained. More specifically, in the case of car insurance claims, this method can obtain the vehicle accident pictures and descriptive text materials of the accident pictures uploaded by the car insurance claim application terminal, so as to use the multimodal data analysis method to conduct vehicle claim review and improve the accuracy of vehicle claim review.

[0024] Step 202: Extract text feature information and image feature information from the information extraction material; Specifically, the extraction of text and image features from the information extraction material can be directly implemented using the text feature extraction network and image feature extraction network of the multimodal recognition large model. The text feature extraction network can be, for example, a bag-of-words model-based text feature extraction network, and the image feature extraction network can include a CNN-based image feature extraction network. The number of extraction layers in the text and image feature extraction networks is determined by the actual deep learning extraction requirements, with each layer extracting feature information from different dimensions. Generally, the more extraction layers in the text and image feature extraction networks, the stronger the deep learning capability.

[0025] In this embodiment, before extracting textual and image features from the information extraction material, the multimodal recognition model is pre-trained using a large amount of relevant knowledge through deep learning. This pre-trained multimodal recognition model is able to extract rich textual and image features from the newly acquired information extraction material. Specifically, in the vehicle claims review scenario, the pre-training of the multimodal recognition model uses a large number of vehicle accident images and corresponding image description texts for each accident image. Ultimately, the pre-trained multimodal recognition model not only learns sufficient claims-related knowledge but also can perform image feature extraction and text feature extraction for newly input vehicle accident images and image description texts, respectively.

[0026] Step 203: Merge the text feature information and the image feature information to obtain fused feature information.

[0027] In this embodiment, the fusion of text feature information and image feature information to obtain fused feature information can be performed using a parallel fusion approach. Specifically, data from different modalities are input into their respective sub-networks for feature extraction, and then the extracted features are fused. Common fusion methods include element-wise addition, concatenation, and weighted summation. This strategy maintains the independence of each modal data while enhancing the model's expressive power through feature fusion. For example, the text feature extraction network and image feature extraction network extract corresponding text features and image features, respectively. Then, element-wise addition, concatenation, and weighted summation are used to fuse the text features and image features to obtain the fused feature. Alternatively, an embedded fusion approach can be used to obtain the fused feature information. Specifically, the text feature information and image feature information obtained in step 202 are mapped to the same low-dimensional vector space for same-dimensional representation. Then, the text feature representation and image feature representation are fused in this low-dimensional vector space to obtain the fused feature.

[0028] Step 204: Extract target visual features and target semantic features from the fused feature information to obtain feature extraction results.

[0029] Specifically, in step 203, after obtaining the low-dimensional fusion features, the fusion feature information can be transformed and mapped to high dimensions according to the feature dimensions corresponding to the target video features and target semantic features, respectively. Then, the target visual features and target semantic features are obtained as the feature extraction results.

[0030] In this embodiment, the target visual features and target semantic features include, for example, the image features and semantic features of the vehicle generated due to the accident, including collision information, scratch information and dent information; and, for example, the vehicle's own body cleanliness features not caused by the accident, such as the dust and stains on the vehicle's surface.

[0031] Step 205: Based on the feature extraction results and the preset feature differentiation strategy, select the desired feature results.

[0032] Specifically, assuming that in step 204, the feature extraction result includes image features and semantic features of the vehicle generated due to an accident, including collision information, scratch information, and dent information, as well as vehicle body cleanliness features inherent to the vehicle not caused by an accident, such as the dust and stains on the vehicle surface, the desired feature in step 205 is the vehicle body cleanliness feature inherent to the vehicle not caused by an accident. In this case, according to a preset feature discrimination strategy, such as filtering information on the dust and stains on the vehicle surface based on feature description keywords, the desired feature result is obtained. Conversely, if the desired feature in step 205 is image features and semantic features of the vehicle generated due to an accident, then according to a preset feature discrimination strategy, such as filtering information on collision information, scratch information, and dent information on the vehicle surface based on feature description keywords, the desired feature result is obtained. It should be understood that the desired feature result is set according to specific business processing requirements.

[0033] Step 206: Based on the expected feature results and the preset level division strategy, determine the representation level of the information extraction material, and perform hierarchical representation of the target object, wherein the information extraction material is the information extraction material of the target object.

[0034] Specifically, the preset level classification strategy, for example, involves setting three levels for the cleanliness of the vehicle itself based on the learning progress of the multimodal recognition model during pre-training, according to the range of feature values: "untidy body," "slightly untidy body," and "clean body." Subsequently, when the desired feature result is obtained, i.e., after acquiring information such as dust and stains on the vehicle surface, the cleanliness level of the vehicle is determined. Alternatively, the preset level classification strategy could also involve setting multiple levels for the vehicle's accident status based on the learning progress of the multimodal recognition model during pre-training, according to the range of feature values: "minor scratches," "minor impacts," and "severe impacts." Subsequently, when the desired feature result is obtained, i.e., after acquiring information such as collision information, scratch information, and dent information on the vehicle surface, the collision level of the vehicle is determined. The target object includes the vehicle's body condition, which includes collision status or cleanliness status, etc. The target object varies depending on the output target.

[0035] The data representation method provided in this embodiment involves: acquiring information extraction materials uploaded by the target data terminal; extracting textual and image feature information from the information extraction materials; fusing the textual and image feature information to obtain fused feature information; extracting target visual and semantic features from the fused feature information to obtain feature extraction results; filtering desired feature results based on the feature extraction results and a preset feature differentiation strategy; determining the representation level of the information extraction materials based on the desired feature results and a preset level division strategy, and performing hierarchical representation of the target object, wherein the information extraction materials are the information extraction materials of the target object. In the context of auto insurance claims in the field of financial technology, this method can perform feature extraction, fusion, and desired feature result filtering on acquired vehicle accident images and descriptive text materials. It employs multimodal data analysis to analyze vehicle claims data and performs hierarchical representation of the data analysis results for the target vehicle, enriching the data analysis methods and improving the accuracy of target data analysis.

[0036] In this embodiment, before performing the step of extracting text feature information and image feature information from the information extraction material, the method further includes: inputting the information extraction material into a pre-trained multimodal recognition large model, wherein the multimodal recognition large model includes a deep learning-based multimodal recognition large model, and the deep learning-based multimodal recognition large model includes an image feature extraction layer, a text feature extraction layer, and a feature fusion layer in its neural network processing layer structure; the image feature extraction layer contains multiple layers of image feature extraction neural networks, and the text feature extraction layer also contains multiple layers of text feature extraction neural networks; the feature fusion layer can be added after the last neural network layer of the image feature extraction layer and the text feature extraction layer, or it can be added alternately between each neural network layer of the image feature extraction layer and the text feature extraction layer. The specific architecture is designed according to the actual processing requirements and is not limited here. The difference is that adding the feature fusion layer after the last neural network layer of the image feature extraction layer and the text feature extraction layer is suitable for parallel fusion; while adding it alternately between each neural network layer of the image feature extraction layer and the text feature extraction layer is more suitable for serial fusion.

[0037] In this embodiment, the step of extracting text feature information and image feature information from the information extraction material includes: using the information extraction component in the pre-trained multimodal recognition large model to extract text feature information and image feature information from the information extraction material, respectively. The information extraction component includes a text feature extraction component and an image feature extraction component, namely a text feature extraction layer and an image feature extraction layer. The text feature extraction layer and the image feature extraction layer are used to extract text feature information and image feature information from the information extraction material, respectively.

[0038] Continue to refer to Figure 3 Before performing step 202, Figure 3 The flowchart shown is a specific embodiment of the data representation method described in this application for pre-training a large multimodal recognition model, including the following steps: Step 301: Obtain model training data, wherein the model training data includes text data and image data used for training; Specifically, in the aforementioned vehicle claims review scenario, the image data used for training includes a large number of vehicle accident images, and the text data used for training includes the image description text corresponding to each vehicle accident image.

[0039] Step 302: Input the model training data into the multimodal recognition large model to be trained; In this embodiment, the large multimodal recognition model to be trained refers to a large multimodal recognition model that has a determined number of feature extraction layers, model hyperparameters, etc., and is to be trained using deep learning.

[0040] Step 303: Using deep learning, learn the target knowledge contained in the text data and image data used for training, respectively; In this embodiment, the target knowledge, for example in a vehicle claims business scenario, includes data knowledge describing the severity of an accident, data knowledge describing claims liability, etc.; it should be understood that the learning of the target knowledge may vary depending on the business scenario or the learning objective.

[0041] Step 304: Based on a preset knowledge differentiation strategy, classify the target knowledge into interference knowledge class and expected knowledge class; Specifically, in conjunction with the above example, the target knowledge includes knowledge data describing the cleanliness of the vehicle itself and knowledge data describing the collision severity of the vehicle body. Here, the target knowledge is classified into interference knowledge class and expected knowledge class. Specifically, the knowledge data describing the cleanliness of the vehicle itself is classified as expected knowledge class, and the knowledge data describing the collision severity of the vehicle body is classified as interference knowledge class. Conversely, the knowledge data describing the cleanliness of the vehicle itself is classified as interference knowledge class, and the knowledge data describing the collision severity of the vehicle body is classified as expected knowledge class.

[0042] Step 305: Based on the target knowledge contained in the expected knowledge class and the preset level setting strategy, set the material evaluation level in the multimodal recognition large model to be trained, and obtain the pre-trained multimodal recognition large model.

[0043] Specifically, assuming the expected knowledge class is knowledge data describing the cleanliness of the vehicle itself; the preset level setting strategy includes setting a material evaluation level based on the description values ​​corresponding to the knowledge data describing the cleanliness of the vehicle itself; subsequently, the description values ​​corresponding to the knowledge data describing the cleanliness of the vehicle itself and the corresponding set material evaluation levels are imported into the multimodal recognition model as recognition knowledge to obtain the pre-trained multimodal recognition model.

[0044] In this embodiment, a deep learning approach is adopted, combined with multimodal feature extraction and fusion methods, to pre-train a large multimodal recognition model. In particular, in the vehicle claims review scenario, this deep learning approach is first used to pre-train a large multimodal recognition model, so that when reviewing vehicle claims, multimodal information can be acquired and represented based on the input information materials, and finally the hierarchical representation result of the target vehicle can be obtained.

[0045] Continue to refer to Figure 4 , Figure 4 yes Figure 2 A flowchart of a specific embodiment of step 204 shown includes the following steps: Step 401: Identify the feature representations corresponding to the target's visual features and semantic features, respectively; Specifically, after acquiring the fused feature information, the fused feature information is transformed and mapped to a higher dimension corresponding to the target visual feature. Then, using a preset visual dimension information recognition template, the fused feature matrix corresponding to the transformed target visual feature is identified. That is, the feature representation of the target visual feature is the fused feature matrix corresponding to the target visual feature. Alternatively, the fused feature information is transformed and mapped to a higher dimension corresponding to the target semantic feature. Then, using a preset semantic dimension information recognition template, the fused feature matrix corresponding to the transformed target semantic feature is identified. That is, the feature representation of the target semantic feature is the fused feature matrix corresponding to the target semantic feature. Here, "feature representation" can be understood as representation data in the form of a feature encoding matrix.

[0046] Step 402: Extract target visual features and target semantic features from the fused feature information according to the feature representation, and obtain target visual feature extraction results and target semantic feature extraction results.

[0047] Specifically, the fused feature information refers to the fused features obtained by performing feature fusion using element-level addition, concatenation, and weighted summation. For example, it is a new feature matrix obtained by concatenating corresponding positions of two feature matrices. In this case, the fused feature information contains at least the feature representations contained in the two feature matrices respectively. Assuming that these two feature matrices are a visual feature matrix and a semantic feature matrix, the fused feature matrix corresponding to the transformed target visual features is used as the basis for extracting target visual features. And, the fused feature matrix corresponding to the transformed target semantic features is used as the basis for extracting target semantic features. Target visual features and target semantic features can be extracted from the fused feature information. Finally, the fused feature matrix is ​​converted into specific text to obtain the target visual feature extraction result and target semantic feature extraction result in text data form. Specifically, the target visual feature extraction result in text data form is obtained based on the visual feature text representation data corresponding to each matrix code value in the fused feature matrix corresponding to the target visual features. Correspondingly, the target semantic feature extraction result in text data form is obtained based on the semantic feature text representation data corresponding to each matrix code value in the fused feature matrix corresponding to the target semantic features.

[0048] Continue to refer to Figure 5 , Figure 5 yes Figure 2 A flowchart of a specific embodiment of step 205 shown includes the following steps: Step 501: Identify the visual feature extraction results and semantic feature extraction results contained in the interference knowledge class from the feature extraction results; Specifically, in this embodiment, based on the actual output requirements of the business scenario, interference knowledge class and expected knowledge class are divided. Here, the visual feature extraction results and semantic feature extraction results contained in the interference knowledge class are identified from the feature extraction results so that only the feature extraction results corresponding to the expected knowledge class are retained in the future.

[0049] Step 502: Delete all visual feature extraction results and all semantic feature extraction results contained in the interference knowledge class in the feature extraction results, and obtain only the remaining feature extraction results as the expected feature results.

[0050] In this embodiment, after performing the step of deleting all visual feature extraction results and all semantic feature extraction results contained in the interfering knowledge class from the feature extraction results, and obtaining only the remaining feature extraction results as the expected feature results, the method further includes: identifying whether the deleted visual feature extraction results and semantic feature extraction results simultaneously belong to the visual feature extraction results and semantic feature extraction results of the expected knowledge class; if the currently deleted visual feature extraction result also belongs to the visual feature extraction results of the expected knowledge class, or the currently deleted semantic feature extraction result also belongs to the semantic feature extraction results of the expected knowledge class, then the currently deleted visual feature extraction result or semantic feature extraction result is added to the expected feature results to update the expected feature results.

[0051] Specifically, since there may be overlap in feature information between multi-dimensional feature dimensions, that is, there are mutual influences and correlations between multi-dimensional features, it is necessary to identify whether the deleted visual feature extraction results and semantic feature extraction results belong to both the visual feature extraction results and semantic feature extraction results of the expected knowledge class at the same time, so as to avoid excessive deletion causing the expected knowledge class to be mistakenly deleted.

[0052] Continue to refer to Figure 6 , Figure 6 yes Figure 2 A flowchart of a specific embodiment of step 206 shown includes the following steps: Step 601: Identify all knowledge information corresponding to the desired feature result; Step 602: Organize all the knowledge information to construct a knowledge information set; Step 603: Determine the representation level of the information extraction material based on the material evaluation level corresponding to the knowledge information set, and perform hierarchical representation of the target object.

[0053] Specifically, the material evaluation level is set in step 305 by establishing a material evaluation level in the multimodal recognition model to be trained based on the target knowledge contained in the expected knowledge class and a preset level setting strategy. The knowledge learned through deep learning is used to perform a hierarchical representation of the target object. For example, after first training a multimodal knowledge model for vehicle claims review, when new vehicle accident images and image description text are input, the pre-trained multimodal knowledge model is used to perform a hierarchical representation of the target vehicle, such as representing the vehicle body cleanliness level and the vehicle surface impact level. Compared to existing methods that use large amounts of business data for annotation and training, using deep learning for pre-training of the business knowledge model can learn richer and more comprehensive business knowledge, and at the same time, it can improve the accuracy of the output of the expected results in subsequent recognition operations.

[0054] The data representation method provided in this embodiment involves: acquiring information extraction materials uploaded by the target data terminal; extracting textual and image feature information from the information extraction materials; fusing the textual and image feature information to obtain fused feature information; extracting target visual and semantic features from the fused feature information to obtain feature extraction results; filtering desired feature results based on the feature extraction results and a preset feature differentiation strategy; determining the representation level of the information extraction materials based on the desired feature results and a preset level division strategy, and performing hierarchical representation of the target object, wherein the information extraction materials are the information extraction materials of the target object. In the context of auto insurance claims in the field of financial technology, this method can perform feature extraction, fusion, and desired feature result filtering on acquired vehicle accident images and descriptive text materials. It employs multimodal data analysis to analyze vehicle claims data and performs hierarchical representation of the data analysis results for the target vehicle, enriching the data analysis methods and improving the accuracy of target data analysis.

[0055] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0056] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0057] The data representation method provided in this embodiment involves: acquiring information extraction materials uploaded by the target data terminal; extracting textual and image feature information from the information extraction materials; fusing the textual and image feature information to obtain fused feature information; extracting target visual and semantic features from the fused feature information to obtain feature extraction results; filtering desired feature results based on the feature extraction results and a preset feature differentiation strategy; determining the representation level of the information extraction materials based on the desired feature results and a preset level division strategy, and performing hierarchical representation of the target object, wherein the information extraction materials are the information extraction materials of the target object. In the context of auto insurance claims in the field of financial technology, this method can perform feature extraction, fusion, and desired feature result filtering on acquired vehicle accident images and descriptive text materials. It employs multimodal data analysis to analyze vehicle claims data and performs hierarchical representation of the data analysis results for the target vehicle, enriching the data analysis methods and improving the accuracy of target data analysis.

[0058] Further reference Figure 7 As a response to the above Figure 2 To implement the method shown, this application provides an embodiment of a data characterization device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0059] like Figure 7 As shown, the data characterization device 700 described in this embodiment includes: an information extraction material acquisition module 701, a feature information extraction module 702, a feature information fusion module 703, a target feature extraction module 704, a desired feature screening module 705, and a hierarchical characterization module 706. Wherein: The information extraction material acquisition module 701 is used to acquire the information extraction material uploaded by the target data terminal, wherein the information extraction material includes text information extraction material and image information extraction material; Feature information extraction module 702 is used to extract text feature information and image feature information from the information extraction material; The feature information fusion module 703 is used to fuse the text feature information and the image feature information to obtain fused feature information; The target feature extraction module 704 is used to extract target visual features and target semantic features from the fused feature information to obtain feature extraction results. The desired feature filtering module 705 is used to filter out desired feature results based on the feature extraction results and a preset feature differentiation strategy. The hierarchical representation module 706 is used to determine the representation level of the information extraction material based on the expected feature results and the preset level division strategy, and to perform hierarchical representation of the target object, wherein the information extraction material is the information extraction material of the target object.

[0060] This application involves: acquiring information extraction materials uploaded by the target data terminal; extracting textual and image features from the information extraction materials; fusing the textual and image features to obtain fused feature information; extracting target visual and semantic features from the fused feature information to obtain feature extraction results; filtering desired feature results based on the feature extraction results and a preset feature differentiation strategy; determining the representation level of the information extraction materials based on the desired feature results and a preset level division strategy, and performing hierarchical representation of the target object, wherein the information extraction materials are the information extraction materials of the target object. In the context of auto insurance claims in the field of financial technology, this method can perform feature extraction, fusion, and desired feature result filtering on acquired vehicle accident images and descriptive text materials. It employs multimodal data analysis to analyze vehicle claims data and performs hierarchical representation of the data analysis results for the target vehicle, enriching data analysis methods and improving the accuracy of target data analysis.

[0061] In this embodiment, the data characterization device 700 further includes a model processing input module, which is used to input the information extraction material into a pre-trained multimodal recognition large model.

[0062] In this embodiment, the feature information extraction module 702 is specifically used to extract text feature information and image feature information from the information extraction material by utilizing the information extraction component in the pre-trained multimodal recognition large model.

[0063] In this embodiment, the data representation device 700 further includes a training data acquisition module, a training data input module, a target knowledge learning module, a target knowledge classification module, and an evaluation level setting module. Wherein: The training data acquisition module is used to acquire model training data, wherein the model training data includes text data and image data used for training; The training data input module is used to input the model training data into the multimodal recognition large model to be trained; The target knowledge learning module is used to learn the target knowledge contained in the text data and image data used for training, respectively, using deep learning methods. The target knowledge classification module is used to classify the target knowledge based on a preset knowledge differentiation strategy, dividing it into interference knowledge class and expected knowledge class; The evaluation level setting module is used to set the material evaluation level in the multimodal recognition large model to be trained according to the target knowledge contained in the expected knowledge class and the preset level setting strategy, so as to obtain the pre-trained multimodal recognition large model.

[0064] In this embodiment, the target feature extraction module 704 includes a feature representation recognition unit and a target feature extraction unit. Wherein: The feature representation recognition unit is used to recognize the feature representations corresponding to the visual features and semantic features of the target, respectively. The target feature extraction unit is used to extract target visual features and target semantic features from the fused feature information according to the feature representation, and obtain target visual feature extraction results and target semantic feature extraction results.

[0065] In this embodiment, the desired feature screening module 705 includes an interference feature identification unit and an interference feature deletion unit. Wherein: An interference feature identification unit is used to identify the visual feature extraction results and semantic feature extraction results contained in the interference knowledge class from the feature extraction results. The interference feature removal unit is used to remove all visual feature extraction results and all semantic feature extraction results contained in the interference knowledge class in the feature extraction results, so as to obtain only the remaining feature extraction results as the expected feature results.

[0066] In this embodiment, the data characterization device 700 further includes a common feature recognition module and a desired feature update module. Wherein: A common feature recognition module is used to identify whether the deleted visual feature extraction results and semantic feature extraction results belong to both the visual feature extraction results and semantic feature extraction results of the expected knowledge class. The expected feature update module is used to update the expected feature result by adding the currently deleted visual feature extraction result or semantic feature extraction result to the expected feature result if the currently deleted visual feature extraction result also belongs to the visual feature extraction result of the expected knowledge class, or the currently deleted semantic feature extraction result also belongs to the semantic feature extraction result of the expected knowledge class.

[0067] In this embodiment, the hierarchical representation module 706 includes a knowledge information recognition unit, a knowledge information organization unit, and a hierarchical representation unit. Wherein: A knowledge information identification unit is used to identify all knowledge information corresponding to the desired feature result; The knowledge information organization unit is used to organize all the knowledge information and construct a knowledge information set; The hierarchical representation unit is used to determine the representation level of the information extraction material based on the material evaluation level corresponding to the knowledge information set, and to perform hierarchical representation of the target object.

[0068] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the methods described above. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).

[0069] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0070] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 8 , Figure 8 This is a basic structural block diagram of the computer device in this embodiment.

[0071] The computer device 8 includes a memory 8a, a processor 8b, and a network interface 8c that are interconnected via a system bus. It should be noted that... Figure 8Only a computer device 8 with component memory 8a, processor 8b, and network interface 8c is shown. However, it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented instead. Those skilled in the art will understand that the computer device described herein is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0072] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.

[0073] The memory 8a includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 8a may be an internal storage unit of the computer device 8, such as the hard disk or memory of the computer device 8. In other embodiments, the memory 8a may also be an external storage device of the computer device 8, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 8. Of course, the memory 8a may also include both the internal storage unit and its external storage device of the computer device 8. In this embodiment, the memory 8a is typically used to store the operating system and various application software installed on the computer device 8, such as computer-readable instructions for a data representation method. In addition, the memory 8a can also be used to temporarily store various types of data that have been output or will be output.

[0074] In some embodiments, the processor 8b may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. The processor 8b is typically used to control the overall operation of the computer device 8. In this embodiment, the processor 8b is used to execute computer-readable instructions stored in the memory 8a or to process data, such as executing computer-readable instructions for the data representation method described above.

[0075] The network interface 8c may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 8 and other electronic devices.

[0076] The computer device proposed in this embodiment belongs to the field of data analysis technology and is applied in the multimodal data review and representation scenario of auto insurance claims. This application obtains information extraction materials uploaded by the target data terminal; extracts textual and image feature information from the information extraction materials; fuses the textual and image feature information to obtain fused feature information; extracts target visual and semantic features from the fused feature information to obtain feature extraction results; filters desired feature results based on the feature extraction results and a preset feature differentiation strategy; and determines the representation level of the information extraction materials based on the desired feature results and a preset level division strategy, and performs hierarchical representation of the target object, wherein the information extraction materials are the information extraction materials of the target object. In the auto insurance claims scenario in the field of financial technology technology, this method can perform feature extraction, fusion, and desired feature result filtering on the acquired vehicle accident images and descriptive text materials of the accident images. It employs multimodal data analysis to analyze vehicle claims data and performs hierarchical representation of the data analysis results for the target vehicle, enriching the data analysis methods and improving the accuracy of target data analysis.

[0077] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by a processor to cause the processor to perform the steps of the data representation method described above.

[0078] The computer-readable storage medium proposed in this embodiment belongs to the field of data analysis technology and is applied to the multimodal data review and representation scenario in auto insurance claims. This application obtains information extraction materials uploaded by the target data terminal; extracts textual and image feature information from the information extraction materials; fuses the textual and image feature information to obtain fused feature information; extracts target visual and semantic features from the fused feature information to obtain feature extraction results; filters desired feature results based on the feature extraction results and a preset feature differentiation strategy; and determines the representation level of the information extraction materials based on the desired feature results and a preset level division strategy, and performs hierarchical representation of the target object, wherein the information extraction materials are the information extraction materials of the target object. In the auto insurance claims scenario in the field of financial technology technology, this method can perform feature extraction, fusion, and desired feature result filtering on the acquired vehicle accident images and descriptive text materials of the accident images. It employs multimodal data analysis to analyze vehicle claims data and performs hierarchical representation of the data analysis results for the target vehicle, enriching the data analysis methods and improving the accuracy of target data analysis.

[0079] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0080] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to make the disclosure of this application more thorough and comprehensive. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application. Software tools or components not belonging to this company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.

[0081] It should be noted that if any AI model software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. All user personal information involved in the embodiments of this application is authorized (with the knowledge and consent) by the relevant parties or fully authorized by all parties, and the executing entity may obtain it through various legal and compliant means. The collection, storage, use, transmission, provision, and disclosure of the information, data, and signals involved all comply with relevant laws and regulations and do not violate public order and good morals.

Claims

1. A data representation method, characterized in that, Includes the following steps: The information extraction materials uploaded by the target data terminal are obtained, wherein the information extraction materials include text information extraction materials and image information extraction materials; Extract textual and image features from the information material; By fusing the text feature information and the image feature information, fused feature information is obtained; The fused feature information is used to extract target visual features and target semantic features to obtain feature extraction results; Based on the feature extraction results and the preset feature discrimination strategy, the desired feature results are selected; Based on the expected feature results and the preset level division strategy, the representation level of the information extraction material is determined, and the target object is represented by a level. The information extraction material is the information extraction material of the target object.

2. The data representation method according to claim 1, characterized in that, Before performing the step of extracting textual and image feature information from the information extraction material, the method further includes: The extracted information is then input into a pre-trained multimodal recognition model. The steps for extracting textual and image feature information from the information material include: Using the information extraction component in the pre-trained multimodal recognition model, textual and image feature information are extracted from the information extraction material.

3. The data representation method according to claim 2, characterized in that, Before performing the step of inputting the extracted information material into the pre-trained multimodal recognition large model, the method further includes: Obtain model training data, wherein the model training data includes text data and image data used for training; The training data of the model is input into the large multimodal recognition model to be trained; Deep learning is used to learn the target knowledge contained in the text and image data used for training. Based on a preset knowledge differentiation strategy, the target knowledge is classified into interference knowledge class and expected knowledge class; Based on the target knowledge contained in the expected knowledge class and the preset level setting strategy, a material evaluation level is set in the multimodal recognition large model to be trained, thereby obtaining a pre-trained multimodal recognition large model.

4. The data representation method according to claim 1, characterized in that, The step of extracting target visual features and target semantic features from the fused feature information to obtain feature extraction results includes: Identify the feature representations corresponding to the visual features and semantic features of the target, respectively; Based on the feature representation, target visual features and target semantic features are extracted from the fused feature information to obtain target visual feature extraction results and target semantic feature extraction results.

5. The data representation method according to claim 3, characterized in that, The step of filtering out the desired feature results based on the feature extraction results and the preset feature discrimination strategy includes: Identify the visual feature extraction results and semantic feature extraction results contained in the interference knowledge class from the feature extraction results; All visual feature extraction results and all semantic feature extraction results contained in the interfering knowledge class are deleted from the feature extraction results, and the remaining feature extraction results are used as the expected feature results.

6. The data representation method according to claim 5, characterized in that, After performing the step of deleting all visual feature extraction results and all semantic feature extraction results contained in the interfering knowledge class in the feature extraction results, and obtaining only the remaining feature extraction results as the desired feature results, the method further includes: Identify whether the deleted visual feature extraction results and semantic feature extraction results simultaneously belong to the visual feature extraction results and semantic feature extraction results of the expected knowledge class. If the currently deleted visual feature extraction result also belongs to the visual feature extraction result of the expected knowledge class, or the currently deleted semantic feature extraction result also belongs to the semantic feature extraction result of the expected knowledge class, then the currently deleted visual feature extraction result or semantic feature extraction result is added to the expected feature result to update the expected feature result.

7. The data representation method according to any one of claims 5 or 6, characterized in that, The step of determining the representation level of the information extraction material based on the expected feature results and a preset level division strategy, and performing hierarchical representation of the target object, includes: Identify all knowledge information corresponding to the desired feature result; All the knowledge information is organized to construct a knowledge information set; Based on the material evaluation level corresponding to the knowledge information set, the representation level of the information extraction material is determined, and the target object is represented in a hierarchical manner.

8. A data characterization device, characterized in that, include: The information extraction material acquisition module is used to acquire the information extraction material uploaded by the target data terminal, wherein the information extraction material includes text information extraction material and image information extraction material; The feature information extraction module is used to extract text feature information and image feature information from the information extraction material; The feature information fusion module is used to fuse the text feature information and the image feature information to obtain fused feature information; The target feature extraction module is used to extract target visual features and target semantic features from the fused feature information to obtain feature extraction results. The desired feature filtering module is used to filter out desired feature results based on the feature extraction results and a preset feature differentiation strategy. The hierarchical representation module is used to determine the representation level of the information extraction material based on the expected feature results and the preset level division strategy, and to perform hierarchical representation of the target object, wherein the information extraction material is the information extraction material of the target object.

9. A computer device, characterized in that, The method includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the data representation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the data representation method as described in any one of claims 1 to 7.