Cross-modal-based claim fraud identification method and device, equipment and medium
Patent Information
- Application Number
- CN202610941637.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-26
- Publication Date
- 2026-09-18
AI Technical Summary
[0005]本发明实施例提供了一种基于跨模态的理赔欺诈识别方法、装置、设备及介质,旨在解决现有理赔的欺诈识别准确率和效率低、成本高的技术问题
[0010] This application provides a cross-modal fraud identification method, device, equipment, and medium. By integrating a cross-modal cross-verification mechanism that combines voiceprint authenticity verification, collision physical energy calculation, and audio-visual spatiotemporal consistency analysis, a three-dimensional intelligent fraud identification method is constructed. Without human intervention, it can automatically identify abnormal features such as missing or forged sound, energy logic conflicts, and asynchronous audio-visual and audio environments in staged fraud. It effectively solves the technical problems of existing systems being insensitive to physical logic disconnection and susceptible to staged fraud, resulting in low accuracy and efficiency in fraud identification. It significantly improves the accuracy and processing efficiency of fraud prevention in claims, and greatly reduces the amount of manual review and claims risk control costs.
Smart Images

Figure CN122779989A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and is applied to the field of financial technology. In particular, it relates to a method, apparatus, device, and medium for identifying claims fraud based on cross-modal methods. Background Technology
[0002] With the rapid development of artificial intelligence and computer vision technology, vehicle claims have achieved damage assessment in seconds. However, the current damage assessment system generally adopts a purely vision-driven method, which is prone to exposing serious risk control loopholes in practical applications. Among them, staged fraud has become the core bottleneck restricting the further popularization of intelligent claims, causing huge economic losses to the industry every year.
[0003] Specifically, existing damage assessment systems can only identify the type, location, and area of surface damage on vehicles through image recognition, but cannot verify the formation logic of the damage or the authenticity of the collision event. Therefore, fraudsters can use existing vehicle damage or artificially created damage to fabricate accident scenes, or upload fake accident videos by taking still photos to apply for compensation. Such fraud scenarios are difficult to effectively identify by relying solely on visual information.
[0004] Therefore, the existing system has technical defects such as a single information modality and insufficient verification dimensions, which directly leads to a low accuracy rate in identifying fraud in claims and a high rate of missed detection of abnormal cases. A large number of cases still need to be manually reviewed, which not only weakens the efficiency advantage of automatic loss assessment, but also increases the operating costs and claims risk control pressure of insurance companies. Summary of the Invention
[0005] This invention provides a method, apparatus, device, and medium for identifying fraud in claims based on cross-modal methods, aiming to solve the technical problems of low accuracy and efficiency and high cost in existing fraud identification methods for claims.
[0006] In a first aspect, embodiments of the present invention provide a cross-modal fraud detection method for claims, comprising: acquiring a claim application uploaded by a user, the claim application including video data of a vehicle accident, accident collision parameters, and geographic data; determining whether the claim application is a fraudulent application based on a preset voiceprint matching model according to the video data, the accident collision parameters, and a preset acoustic library, and outputting a first determination result; determining whether the video data matches the accident collision parameters based on a preset image matching model, determining whether the claim application is a fraudulent application based on the matching result, and outputting a second determination result; calculating a consistency score based on the video data and the geographic data based on a preset scoring calculation model, determining whether the claim application is a fraudulent application based on the consistency score and a preset scoring threshold, and outputting a third determination result; if any of the first determination result, the second determination result, and the third determination result indicates that the claim application is a fraudulent application, then the claim application is marked as a fraudulent application.
[0007] Secondly, embodiments of the present invention also provide a cross-modal claims fraud identification device, which includes a unit for performing the above-described method.
[0008] Thirdly, embodiments of the present invention also provide a computer device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described method.
[0009] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing a computer program, the computer program including program instructions that, when executed by a processor, can implement the above-described method.
[0010] This application provides a cross-modal fraud identification method, device, equipment, and medium. By integrating a cross-modal cross-verification mechanism that combines voiceprint authenticity verification, collision physical energy calculation, and audio-visual spatiotemporal consistency analysis, a three-dimensional intelligent fraud identification method is constructed. Without human intervention, it can automatically identify abnormal features such as missing or forged sound, energy logic conflicts, and asynchronous audio-visual and audio environments in staged fraud. It effectively solves the technical problems of existing systems being insensitive to physical logic disconnection and susceptible to staged fraud, resulting in low accuracy and efficiency in fraud identification. It significantly improves the accuracy and processing efficiency of fraud prevention in claims, and greatly reduces the amount of manual review and claims risk control costs. Attached Figure Description
[0011] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 A schematic flowchart illustrating the cross-modal claims fraud identification method provided in this embodiment of the invention; Figure 2 A schematic diagram of a sub-process of the cross-modal claims fraud identification method provided in an embodiment of the present invention; Figure 3 A schematic diagram of a sub-process of the cross-modal claims fraud identification method provided in an embodiment of the present invention; Figure 4 A schematic diagram of a sub-process of the cross-modal claims fraud identification method provided in an embodiment of the present invention; Figure 5 A schematic diagram of a sub-process of the cross-modal claims fraud identification method provided in an embodiment of the present invention; Figure 6 A schematic diagram of a sub-process of the cross-modal claims fraud identification method provided in an embodiment of the present invention; Figure 7 A schematic diagram of a sub-process of the cross-modal claims fraud identification method provided in an embodiment of the present invention; Figure 8 A schematic block diagram of a cross-modal claims fraud identification device provided in an embodiment of the present invention; Figure 9 A schematic block diagram of a computer device provided for an embodiment of the present invention. Detailed Implementation
[0013] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0014] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0015] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0016] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0017] Please see Figure 1 This is a schematic flowchart illustrating the cross-modal fraud detection method for claims provided in this invention. In this application, the cross-modal fraud detection method for claims is applied to the field of artificial intelligence, particularly in scenarios such as intelligent claims processing within the fintech sector. For example, when a user uploads a video of a vehicle accident for online claims processing via software, the system automatically performs multi-dimensional cross-modal verification on the claim application, identifying fraudulent behaviors such as staged photos, re-reporting old damage, and mismatches between energy and deformation, significantly improving risk control accuracy while maintaining second-level damage assessment efficiency.
[0018] This application provides a method, apparatus, computer device, and storage medium for cross-modal fraud detection in claims. The method includes: acquiring a claim application uploaded by a user, the claim application including video data of a vehicle accident, accident collision parameters, and geographic data; determining whether the claim application is fraudulent based on the video data, the accident collision parameters, and a preset acoustic library using a preset voiceprint matching model, and outputting a first determination result; determining whether the video data matches the accident collision parameters using a preset image matching model, determining whether the claim application is fraudulent based on the matching result, and outputting a second determination result; calculating a consistency score based on the video data and the geographic data using a preset scoring calculation model, determining whether the claim application is fraudulent based on the consistency score and a preset scoring threshold, and outputting a third determination result; if any of the first determination result, the second determination result, and the third determination result indicates that the claim application is fraudulent, then the claim application is marked as fraudulent.
[0019] This application constructs a three-dimensional intelligent fraud detection method for claims by constructing a cross-modal cross-validation mechanism that integrates three dimensions: cross-modal fusion of voiceprint authenticity verification, collision physical energy calculation, and audio-visual spatiotemporal consistency analysis. Without human intervention, it can automatically identify abnormal features such as missing or forged sound, energy logic conflicts, and asynchronous audio-visual and audio environments in staged fraud. It effectively solves the technical problems of existing systems being insensitive to physical logic disconnection and susceptible to staged fraud, resulting in low accuracy and efficiency in fraud detection. It significantly improves the accuracy and processing efficiency of fraud detection for claims, and greatly reduces the amount of manual review, thereby reducing the cost of claims risk control.
[0020] Figure 1 This is a flowchart illustrating the cross-modal claims fraud identification method provided in an embodiment of the present invention. Figure 1 As shown, the method includes the following steps S10-S50.
[0021] S10. Obtain the claim application uploaded by the user, wherein the claim application includes video data of the vehicle accident, accident collision parameters and geographical data; Specifically, the claim application is an electronic data packet submitted by the user through a terminal device (such as a smartphone, computer, etc.). The claim application includes three parts of data: video data of the vehicle accident, accident collision parameters, and geographical data.
[0022] The video data refers to vehicle accident scene videos taken by users through mobile devices, which include a continuous time-series image frame sequence and a synchronous audio stream. The video data should cover the overall appearance of the accident vehicle, close-ups of the damaged parts, and the surrounding environment, and fully record the moment the accident occurred and the damage to the vehicle.
[0023] The collision parameters are structured information that users manually fill in on the claims application page or that is automatically collected by terminal sensors, including but not limited to vehicle model, license plate number, vehicle material parameters, reported collision speed, type of collision object, time and location of the accident.
[0024] The geographic data refers to the location information of the accident, which is automatically collected through the GPS module of the mobile device or the base station positioning. It includes coordinates and a description of the geographic environment (such as inside a tunnel, an open road, or an underground parking garage) obtained by looking up the coordinates.
[0025] This embodiment, upon receiving the claim application data submitted by the user, simultaneously acquires three types of data: video, user-entered information, and geographical location. This provides a unified data foundation for subsequent multi-dimensional cross-validation, eliminating the need for additional user operations and improving the user experience of the claim application process.
[0026] More specifically, after receiving a claim application submitted by a user, the system first performs standardized preprocessing and timestamp alignment on all data. For example, it decodes video data to generate a sequence of video frames arranged in chronological order and a synchronized audio stream; it performs format verification and structured storage on accident collision parameters; and it performs reverse geocoding on geographic data to convert it into structured information containing road type and environmental features, etc., to ensure the accuracy of subsequent cross-modal feature comparison.
[0027] S20. Based on the preset voiceprint matching model, determine whether the claim application is a fraudulent application according to the video data, the accident collision parameters and the preset acoustic library, and output the first judgment result; Specifically, the preset voiceprint matching model is a pre-trained deep neural network model used to extract collision sound features from the audio and compare them with standard collision acoustic features. The preset acoustic library is a standardized database constructed by collecting a large amount of real vehicle collision test data, containing standard collision acoustic features corresponding to different vehicle models, different body materials, different collision speeds, and different collision objects. The first judgment result is the result data of whether the claim application is fraudulent, determined from the voiceprint matching dimension based on the voiceprint matching model, the video data, the accident collision parameters, and the preset acoustic library.
[0028] In practice, the voiceprint matching model determines the fraud type of the claim based on the video data, the accident collision parameters, and a preset acoustic library. Specifically, it compares the audio data in the video data with the vehicle parameters and the acoustic features corresponding to the collision parameters to determine whether the claim is fraudulent or not. If the voiceprint dimensions of the video data and the accident collision parameters are different, the claim is deemed to have a high fraud risk, marked as fraudulent, and a first judgment result indicating fraud is output for subsequent manual review. If the voiceprint dimensions of the video data and the accident collision parameters are the same, the claim is deemed to have a low fraud risk, marked as non-fraudulent, and a first judgment result indicating non-fraud is output.
[0029] The voiceprint matching model described in this embodiment performs fraud identification based on the acoustic features of the video data. By introducing a verification mechanism at the voiceprint dimension, it effectively compensates for the deficiency of existing damage assessment systems in being unable to detect abnormal sound logic, thereby improving the accuracy of fraud identification.
[0030] In one embodiment, reference is made to Figure 2 Step S20 includes steps S21-S24.
[0031] S21. The video data is separated to generate audio data, and the voiceprint features of the audio data are extracted. S22. Match the collision acoustic features of the corresponding vehicle from the preset acoustic library according to the accident collision parameters; S23. Compare the voiceprint features with the collision acoustic features; S24. If the voiceprint feature and the collision acoustic feature do not match, the claim application is determined to be a fraudulent application, and the first judgment result that the claim application is a fraudulent application is output.
[0032] Specifically, the audio data is the sound information recorded at the accident scene separated from the video data, i.e., the audio signal; the voiceprint feature refers to a set of physical acoustic parameters extracted from the audio signal generated by the collision event (such as metal impact, plastic breakage, glass shattering, etc.) to characterize the unique attributes of this collision, which is different from the human voiceprint used for identification. The voiceprint feature in this embodiment mainly includes, but is not limited to: the frequency peak at the moment of collision and the sound energy decay time curve.
[0033] The collision acoustic features refer to the baseline voiceprint feature vectors stored in the preset acoustic library, which are adapted to the specific standard working conditions (in this embodiment, specific vehicle model, specific vehicle material, specific obstacle material, and specific vehicle speed) corresponding to the claim application, and are used to compare with the voiceprint features of the audio data collected on site.
[0034] The fraudulent claim refers to a claim application that carries the risk of fraud.
[0035] In the specific implementation process, after receiving the claim application uploaded by the user, the system first processes the video file contained in the claim application, separates the composite video data into independent video streams and audio streams, extracts the audio track and decodes it into the original audio data format that can be used for subsequent analysis and processing.
[0036] Subsequently, the voiceprint matching model detects collision sound events in the original audio data, locates time segments with abrupt energy changes in the audio, and extracts voiceprint features reflecting the physical properties of the collision from these segments. Simultaneously, the system parses the vehicle model, vehicle material, collision speed, and collision substance type fields entered in the accident collision parameters, constructs a database retrieval statement using these fields as query conditions, and initiates a matching request to a preset acoustic library. The preset acoustic library returns standard collision acoustic feature data corresponding to the claim application.
[0037] The system inputs the voiceprint features extracted from the scene and the collision acoustic features into the similarity calculation module, and uses algorithms such as cosine similarity as a metric to calculate the alignment degree of the two vectors in the feature space. If the calculated similarity is lower than the preset voiceprint matching threshold, it indicates that there is a substantial difference between the voiceprint characteristics of the scene audio and the standard collision acoustic characteristics that should be produced under the collision conditions claimed by the user. Based on this, the system determines that the voiceprint features and collision acoustic features do not match, and then outputs the first judgment result that the claim application is a fraudulent application for subsequent manual review. The specific indicators of voiceprint mismatch (such as similarity score, deviation description, etc.) are attached to the damage assessment report to provide irrefutable evidence, improve the professionalism and scientific nature of the claims process, and reduce the dispute rate between customers.
[0038] For example, if the video data in a claim only contains the sound of wind and human voices and lacks instantaneous high-energy sound wave pulses, it is determined to be a static staged shot, and the claim is marked as a fraudulent claim.
[0039] If the calculated similarity is equal to or higher than the preset voiceprint matching threshold, it indicates that the voiceprint characteristics of the on-site audio are consistent with the standard collision acoustic characteristics that should be generated under the collision conditions claimed by the user. Based on this, the system determines that the voiceprint features match the collision acoustic features, and then outputs the first judgment result that the claim application is a non-fraudulent application.
[0040] This embodiment incorporates auditory information, which is completely missing in traditional systems, into the anti-fraud review process through the aforementioned voiceprint matching mechanism. This enables the system to identify staged fraud techniques such as videos without real collision sounds or fake collision sounds. For example, if a fraudster uses existing vehicle damage for a static staged photo, the video and audio tracks will inevitably lack real impact sound pulses that match the collision energy. Alternatively, the artificially created knocking sounds will deviate significantly from the acoustic characteristics of a real collision in terms of peak frequency and decay time. The voiceprint matching model can accurately capture such acoustic anomalies, effectively compensating for the inherent deficiency of existing systems in their insensitivity to missing sound evidence chains, thereby improving the accuracy of fraud identification.
[0041] In one embodiment, reference is made to Figure 3 The voiceprint features include frequency peaks and decay time curves, and step S21 includes steps S211-S213.
[0042] S211. Perform multi-source noise filtering on the audio data to obtain filtered audio; S212. Convert the filtered audio into a frequency domain spectrum; S213. Identify the frequency peak and decay time curve of the frequency domain spectrum.
[0043] Specifically, the filtered audio refers to the pure audio data obtained after filtering from multiple sources of noise, retaining only the effective collision sound signal, including impact sound, cracking sound, vehicle resonance sound, etc. The frequency domain spectrum refers to the signal distribution obtained after converting the time-domain audio signal into a frequency domain representation, used to analyze the frequency composition of the sound. The frequency peak refers to the frequency point in the frequency domain spectrum where the energy amplitude reaches its maximum value. The frequency peak is closely related to the material contact stiffness and natural vibration frequency of the colliding object and is a key acoustic indicator for distinguishing different collision object materials. The decay time curve refers to the trajectory of the collision sound energy decaying from the peak level to the background noise level over time. The decay time curve reflects the damping characteristics and energy dissipation rate of the collision system; collisions with different material combinations exhibit significantly different decay patterns.
[0044] In practice, the system extracts the audio track from user-uploaded videos and decodes it to obtain the raw audio data. Then, it performs multi-source noise filtering on the audio data to remove various environmental interference signals, resulting in filtered audio. Next, it performs a Fast Fourier Transform on the filtered audio, performing frame-by-frame calculations after windowing to generate a frequency domain spectrogram with time and frequency as the coordinate axes. On the spectrogram, the system searches along the frequency dimension for the frequency coordinate with the highest energy amplitude within the collision event time window as the frequency peak; simultaneously, it tracks the energy change sequence of the frequency band where the frequency peak is located along the time dimension, extracting the envelope trajectory of energy decaying from the peak to the background noise level, i.e., the decay time curve. Both together constitute the voiceprint feature vector.
[0045] When comparing voiceprint features and collision acoustic features, the system will perform similarity matching between the spectral features (i.e., frequency peaks, such as energy distribution and impact duration) and decay time curves of real-time audio and the corresponding physical models in the acoustic library.
[0046] This embodiment effectively suppresses complex environmental noise through multi-source noise filtering, quantifies the physical properties of collision sound through fast Fourier transform, and characterizes the collision acoustic features from both frequency and time domains using peak frequency and decay time curves, providing a stable and quantifiable basis for subsequent voiceprint matching.
[0047] S30. Based on a preset image matching model, determine whether the video data matches the accident collision parameters, determine whether the claim application is a fraudulent application based on the matching result, and output a second judgment result; Specifically, the preset image matching model refers to a pre-constructed computational model used to reconstruct a collision damage model based on a video image frame sequence and verify its energy logic consistency with the collision conditions (such as collision speed and collision severity) in the accident collision parameters. The image matching model integrates computer vision 3D reconstruction technology with collision mechanics energy calculation algorithms, enabling it to verify the logical matching between collision energy input and deformation damage output from a physical perspective. The second judgment result is the result data determining whether the claim application is fraudulent based on the image matching model, the video data, and the accident collision parameters from the image energy logic matching dimension.
[0048] In the specific implementation process, the image matching model models vehicle damage based on the image frame data of the video data, and then logically matches the modeling data with the collision parameters in the accident collision parameters. If the two do not match, the claim application is determined to have a high fraud risk, the claim application is marked as a fraudulent application, and the second judgment result of the claim application being a fraudulent application is output for subsequent manual review; if the two match, the claim application is determined to have a low fraud risk, the claim application is marked as a non-fraudulent application, and the second judgment result of the claim application being a non-fraudulent application is output.
[0049] This embodiment introduces a collision mechanics energy calculation mechanism, enabling the system to possess the physical reasoning ability to determine whether energy and deformation match. For example, in staged photo fraud scenarios, fraudsters often use existing vehicle damage to fabricate claims or artificially create minor scratches to exaggerate collision speeds. Such behavior inevitably leads to a significant order-of-magnitude difference between the claimed input energy and the actual energy absorbed by the damage. The image matching model can effectively identify such physical logic conflicts, making up for the shortcomings of existing systems in identifying energy logic contradictions and improving the accuracy of fraud detection.
[0050] In one embodiment, reference is made to Figure 4 The collision parameters include vehicle parameters, collision speed and collision material, and step S30 includes steps S31-S35.
[0051] S31. The video data is separated to generate image frame data; S32. Perform damage modeling and calculate the absorbed energy during the collision based on the image frame data and the vehicle parameters; S33. Calculate the theoretical kinetic energy reference value of the colliding vehicle speed and the colliding material, and determine whether the theoretical kinetic energy reference value matches the absorbed energy; S34. If a match is found, the claim application is determined to be a non-fraudulent application, and the second determination result that the claim application is a non-fraudulent application is output. S35. If there is no match, the claim application is determined to be a fraudulent application, and the second determination result that the claim application is a fraudulent application is output.
[0052] Specifically, the image frame data refers to a sequence of high-definition images (multiple frames) extracted from user-uploaded video data, covering different perspectives of the vehicle damage area. To meet the requirements of subsequent 3D damage modeling for multi-view geometric constraints, the image sequence must have sufficient parallax coverage and inter-frame overlap. Damage modeling refers to the process of 3D reconstruction of the surface geometry of the vehicle damage area based on the multi-view image frame data. Its output is a digital surface model of the damage area, which can be used to quantitatively calculate the indentation depth, deformation area, and plastic deformation volume of the vehicle damage. Absorbed energy refers to the energy dissipated by the plastic deformation of the vehicle body panel material during a collision. Its value is directly related to the indentation volume of the damage area, the yield strength of the body material, and the strain rate correction coefficient.
[0053] In the specific implementation process, the system first separates and processes the video data uploaded by the user, extracts the video track, and filters keyframes according to a preset sampling strategy to generate image frame data covering the damaged area. Subsequently, the three-dimensional reconstruction module in the image matching model receives the image frame data as input, reconstructs the damage model through a preset modeling algorithm, and calculates the damage parameters of the damaged area relative to the undamaged reference surface based on the damage model. In this embodiment, the damage parameters include the length, depth, and stress area of the vehicle damage. The system also reads the material parameters of the vehicle body panel from the vehicle parameter field in the accident collision parameters and calculates the absorbed energy required to generate the volumetric deformation according to solid mechanics formulas.
[0054] Meanwhile, the image matching model reads two fields from the accident collision parameters: collision speed and collision material. Based on the collision material type, it determines the equivalent mass or stiffness characteristics of the collision object and calculates the theoretical kinetic energy reference value corresponding to the accident condition, combined with the collision speed. This theoretical kinetic energy reference value refers to the collision kinetic energy value corresponding to the accident collision parameters described in the user's submitted claim. The image matching model logically compares the calculated absorbed energy with the theoretical kinetic energy reference value. If the difference between the absorbed energy and the theoretical kinetic energy reference value falls within a preset reasonable range, it determines that the absorbed energy logically matches the energy of the collision speed and collision material, and outputs the conclusion that the claim is a non-fraudulent claim, i.e., outputting the second judgment result that the claim is a non-fraudulent claim. If the absorbed energy is significantly lower than the theoretical kinetic energy reference value, i.e., the difference between the absorbed energy and the theoretical kinetic energy reference value is lower than the lower limit of the preset reasonable range, it indicates that the actual damage is far less than the deformation that should have been caused by the claimed impact, and energy is present. The logic of large losses and small losses is contradictory. If the absorbed energy is significantly higher than the theoretical kinetic energy reference value, that is, the difference between the absorbed energy and the theoretical kinetic energy reference value is higher than the upper limit of the preset reasonable range, it indicates that the actual damage is abnormally severe, which may be due to the accumulation of multiple collisions or the superposition of old injuries. Both of the above mismatch situations are judged as energy logic mismatch, and a fraudulent claim conclusion is output. That is, the second judgment result that the claim application is a fraudulent claim is output. The conclusion of the above mismatch can be attached to the loss assessment report to provide irrefutable evidence, improve the professionalism and scientific nature of the claim, and reduce the dispute rate between customers.
[0055] For example, if the vehicle damage in the video data of a claim requires a huge kinetic energy after calculation, but the user describes it as a low-speed light collision, then the two capabilities do not match, and the claim will be marked as a fraudulent claim.
[0056] This embodiment combines multi-view 3D reconstruction technology with collision mechanics energy calculation algorithms, enabling the system to reason about whether energy input and deformation output match from a physical perspective. For example, in staged photo fraud scenarios, fraudsters often use existing vehicle damage to falsely report or exaggerate collision speeds. Such behavior inevitably leads to an order-of-magnitude difference between the claimed input energy and the actual energy absorbed by the damage. Image matching models can effectively identify such physical logic conflicts, compensating for the inherent deficiency of existing systems in being unable to identify contradictions in energy logic and improving the accuracy of fraud detection.
[0057] Meanwhile, damage assessment reports based on energy calculations provide irrefutable physical evidence for claim denials, reducing disputes with customers and enhancing the professionalism and scientific rigor of claims processing.
[0058] More specifically, the step of performing damage modeling and calculating the absorbed energy during a collision based on the image frame data and the vehicle parameters includes: extracting the length, depth, and force-bearing area of the vehicle damage region from the image frame data; and calculating the absorbed energy based on the length, depth, force-bearing area, and the vehicle parameters.
[0059] Specifically, the length of the damaged area refers to the maximum linear dimension of the surface of the damaged vehicle component along the main deformation direction; the depth refers to the maximum normal indentation distance of the damaged surface relative to the original undamaged contour; and the force-bearing area refers to the projected area of the damaged area on a plane perpendicular to the direction of the impact force.
[0060] In practical implementation, the system automatically measures three geometric parameters of the damaged area—length, depth, and stress area—based on an image matching model and multi-view image frame data covering the damaged area. Subsequently, the system reads the material parameters of the damaged component from the vehicle parameter field of the accident collision parameters and retrieves the corresponding yield strength value from a preset material database. Based on the length, depth, stress area, and yield strength, the system calculates the absorbed energy, equating the damaged area to a prismatic depression with the stated depth and stress area. The absorbed energy is obtained by adjusting the product of the yield strength and the deformation volume using a dynamic correction coefficient.
[0061] This embodiment achieves a precise quantitative description of collision damage by automatically extracting damage geometric parameters from multi-view image frames. Combined with material parameters, it calculates the absorbed energy, providing an objective and quantifiable basis for subsequent logical comparison with the theoretical kinetic energy corresponding to the user's claimed collision conditions.
[0062] S40. Based on the preset scoring calculation model, calculate the consistency score according to the video data and the geographic data, and determine whether the claim application is a fraudulent application based on the consistency score and the preset scoring threshold, and output the third judgment result.
[0063] Specifically, the preset scoring calculation model refers to a calculation model used to comprehensively evaluate the audio-visual synchronization and environmental acoustic consistency, and output a quantitative risk score. The consistency score is a scalar value that comprehensively reflects the degree of audio-visual synchronization within the video data and the degree of matching between audio features and the acoustic characteristics of the geographical environment. The higher the consistency score, the better the audio-visual data, audio, and geographical environment match, i.e., the lower the fraud risk. The preset scoring threshold is a pre-set threshold obtained through multiple experiments and can be adjusted according to different scenarios. The third judgment result is the result data on whether the claim application is fraudulent, determined based on the scoring calculation model, video data, and geographical data from the audio-visual spatiotemporal consistency score dimension.
[0064] This embodiment uses the scoring calculation model to perform a consistency score on the video data and geographic data, thereby identifying fraud in the claim application from the dimensions of video data and geographic environment information. If the consistency score is lower than a preset scoring threshold, the claim application is determined to be a fraudulent application, and the third judgment result of the claim application being a fraudulent application is output so that the claim application can be manually reviewed; if the consistency score is higher than or equal to the preset scoring threshold, the claim application is determined to be a non-fraudulent application, and the first judgment result of the claim application being a non-fraudulent application is output.
[0065] This embodiment uses a scoring calculation model to identify fraud in claims applications from the dimensions of video data and geographic environment information, achieving a balance between efficiency and risk control. This optimizes resource allocation in the claims process, improves claims efficiency, and reduces costs.
[0066] In one embodiment, reference is made to Figure 5 The consistency score includes audio-visual matching score and environmental matching score, and step S40 includes steps S41-S44.
[0067] S41. The video data is separated to generate audio data and image frame data; S42. Calculate the audio-visual matching score based on the audio data and the image frame data; S43. Calculate the environmental matching score between the audio data and the geographic data; S44. The audio-visual matching score and the environmental matching score are weighted to calculate the consistency score.
[0068] Specifically, the audio-visual matching score is a quantitative score that measures the synchronization between the timing of the collision sound pulse and the timing of the vehicle shaking in the video. A higher score indicates better audio-visual synchronization. The environmental matching score is a quantitative score that measures the consistency between the reverberation characteristics of the on-site recording and the expected acoustic characteristics of the environment type indicated by geographical data. For example, open roads have short reverberation times, while enclosed garages have significantly longer reverberation times. The weighting refers to the summation of the two scores after assigning preset weights to each, with the weights determined through statistical optimization based on their respective discriminative power in fraud detection.
[0069] In practice, the system separates video data to generate audio data and image frame data. The audio-visual matching score is calculated using a millisecond-level timestamp alignment algorithm, based on the audio data and image frame data. Then, an environmental matching score is calculated based on the audio data and geographic data. Finally, the audio-visual matching score and the environmental matching score are weighted and summed according to preset weights to output a consistency score, which is then compared with a preset score threshold.
[0070] Since it is difficult to achieve millisecond-level physical alignment in manually synthesized videos, this embodiment verifies the authenticity of the collision event by calculating the audio-visual matching score from the temporal dimension of audio and image frames, thereby performing audio-visual synchronization verification on the video data; then, it verifies the consistency of the recording location (i.e., environmental fingerprint alignment) by calculating the environmental matching score from the spatial dimension of audio and geographical environment, in order to determine whether the claim application is a secondary on-site fraud, that is, to determine whether the fraudster moved the vehicle to another fake scene after the damage occurred elsewhere to commit fraud; this embodiment uses audio-visual synchronization verification and environmental fingerprint alignment to complement each other for comprehensive evaluation, effectively improving the coverage and accuracy of fraud identification.
[0071] In one embodiment, reference is made to Figure 6 Step S42 includes steps S421-S424.
[0072] S421. Extract the time points of the collision sound pulses based on the audio data; S422. Extract the starting point of vehicle displacement or shaking based on the image frame data; S423. Calculate the time offset between the time point and the starting point based on a preset timestamp alignment algorithm; S424. Calculate the audio-visual matching score based on the time offset.
[0073] Specifically, the time point of the collision sound pulse refers to the precise moment corresponding to the instantaneous energy change caused by the collision event in the audio waveform. Since the collision sound is a typical transient broadband pulse signal, its time-domain waveform exhibits a steep rising edge, which can be accurately located on the time axis through an energy change detection algorithm. The starting point of vehicle displacement or shaking refers to the timestamp corresponding to the starting frame in the video frame where the vehicle changes from a stationary state to a moving state due to external impact, or produces an instantaneous vibration that is visible to the naked eye. In real collision events, the generation of the collision sound pulse and the onset of vehicle shaking occur almost simultaneously in physical terms, and the time difference between the two is only affected by the sound propagation delay and the sampling synchronization error of the recording equipment. The time offset refers to the absolute time difference between the collision sound pulse time point and the starting point of vehicle displacement or shaking. The audio-visual matching score is a score used to quantify the degree of audio-visual synchronization, obtained by converting the time offset through a preset mapping function.
[0074] In the specific implementation process, the system first extracts the time point of the collision sound pulse from the audio data based on the preset audio extraction algorithm. For example, the time point can be extracted by a dual-threshold endpoint detection algorithm based on short-time energy and zero-crossing rate. Then, it extracts the starting point of vehicle movement or shaking from the image frame data based on the preset image frame extraction algorithm. For example, the starting point can be extracted by a motion estimation network based on dense optical flow.
[0075] The absolute difference between the timestamps of the collision sound pulse and the starting point of the vehicle displacement or shaking is calculated using a millisecond-level timestamp alignment algorithm as the time offset between the two. Finally, the time offset is converted into an audio-visual matching score through linear mapping. The smaller the time offset, the higher the audio-visual matching score, indicating that the audio at the time of the vehicle collision and the time of the vehicle collision are closer, and the lower the risk of fraud in the video data.
[0076] This embodiment can effectively identify staged fraudulent activities such as post-synthesized collision sounds and asynchronous audio-visual presentations by accurately verifying the temporal consistency between collision sounds and vehicle physical movements. It fills the gap in existing technologies that cannot verify the physical logic of audio-visual presentations, complements the verification of voiceprint authenticity, and further enhances the fraud detection capability of intelligent claims processing.
[0077] In one embodiment, reference is made to Figure 7 Step S43 includes steps S431-S432.
[0078] S431. Extract the reverberation features of the audio data; S432. Calculate the environmental matching score of the reverberation feature and the geographic data.
[0079] Specifically, the reverberation characteristics refer to a set of acoustic parameters reflecting the size of the recording environment, boundary reflection characteristics, and acoustic attenuation patterns. These mainly include reverberation time in different frequency bands, early reflection arrival delay, and direct-to-reverberation energy ratio. The reverberation characteristics are determined by the physical geometry of the recording location. For example, open roads lack lateral reflecting surfaces and exhibit extremely short reverberation times, while underground parking lots or indoor garages have longer reverberation times and denser early reflections due to multiple reflections from concrete walls. The environmental matching score is a quantitative similarity score calculated by comparing the reverberation characteristics extracted from the on-site recording with the standard reverberation characteristics corresponding to the environmental type indicated by geographical data. A higher score indicates a stronger consistency between the acoustic environment of the recording location and the claimed location.
[0080] In the specific implementation process, after the system completes the separation and generation of audio data, it inputs the audio data into a pre-constructed reverberation feature extraction network. For example, the reverberation feature extraction network adopts a blind reverberation estimation model based on a deep convolutional recurrent architecture. The model's input is the original audio time-domain waveform or a log-Mel spectrum. The model first extracts the time-frequency representation of the audio signal through a multi-layer convolutional neural network, then captures the long-term temporal dependencies in the reverberation decay process through a bidirectional long short-term memory network, and finally outputs a reverberation feature vector through a fully connected regression layer. Each dimension of the reverberation feature vector encodes acoustic parameters such as the reverberation time value within different octave bands of the center frequency, the energy ratio of early reflected sound to direct sound, and the arrival time distribution of reflected sound. Simultaneously, based on the geographical data attached to the user's uploaded claim application, the system extracts the latitude and longitude coordinate information, calls a preset map service interface for inverse geocoding, and obtains the location category label corresponding to the latitude and longitude coordinate information, such as "open road," "in tunnel," and "underground parking garage." The system uses the location category label as a query index to retrieve the corresponding standard reverberation feature vector from a preset environmental acoustic template library. The preset environmental acoustic template library is pre-built offline by collecting impulse response data in different typical environments and extracting reverberation feature parameters. After obtaining the on-site reverberation feature vector and the standard reverberation feature vector, the system calculates the similarity between the two vectors (for example, using a cosine similarity algorithm) and converts the similarity value into an environmental matching score through a linear mapping. The higher the similarity, the higher the environmental matching score, indicating that the audio at the time of the vehicle collision and the location of the vehicle collision are more consistent, and the lower the risk of fraud in the claim application. Conversely, it indicates that the claim application has the risk of secondary on-site fraud.
[0081] This step automatically verifies the acoustic consistency between the recording location and the claimed location by checking the matching of the on-site audio reverberation characteristics with the geographical environment of the reported incident. It can effectively identify fraudulent activities such as staged photos taken in different locations and indoor fake outdoor accidents, filling the gap in existing technologies that cannot verify the authenticity of the accident location. At the same time, it complements voiceprint verification, logic verification, and audio-visual synchronization verification in multiple dimensions, further improving the cross-modal anti-fraud system and enhancing the rigor of risk control in intelligent claims processing.
[0082] For example, in the review of car insurance claims, fraudsters staged an accident in an underground parking garage using old vehicle damage, but reported that the accident occurred on an open main road in the city. The audio reverberation features extracted by this module are seriously inconsistent with the standard acoustic features of open roads, and the calculated environmental matching score is significantly low, which provides a key basis for the subsequent determination that the claim application is a fraudulent application.
[0083] S50. If any one of the first judgment result, the second judgment result, and the third judgment result indicates that the claim application is a fraudulent application, then the claim application is marked as a fraudulent application.
[0084] Specifically, after completing verification across three dimensions—voiceprint matching, image energy logical matching, and consistency scoring—the system summarizes the first, second, and third judgment results. The system performs a logical OR operation on the three judgment results. If any judgment result indicates a fraudulent application, the case status of the claim application is marked as a fraudulent application. If the claim application is marked as a fraudulent application, it is transferred to manual review.
[0085] Specifically, the fraudulent claim refers to a claim application that is deemed abnormal by any module of the voiceprint matching model, image matching model, or scoring calculation model, and whose overall risk level exceeds a preset threshold; that is, a claim application with a high fraud risk level. The manual review refers to the process of automatically marking suspicious cases as fraudulent claims by the system to a back-end review system operated by professional claims reviewers, whereby the cases are manually inspected on-site and reviewed via video, and a final claims decision is made.
[0086] In the specific implementation process, after completing the verification of the aforementioned three dimensions—voiceprint matching, image energy logical matching, and consistency scoring—the system summarizes the judgment results of each module. If any module outputs a fraud claim judgment conclusion, the system marks the case status of the claim as a fraud claim. Subsequently, the system automatically generates a structured risk warning report, which includes basic case information, the specific module identifier that triggered the warning, details of abnormal indicators, and corresponding confidence scores, such as the degree of deviation of the frequency peak recorded by the voiceprint matching module, the difference between the energy ratio recorded by the image matching module and the reasonable range, and the difference between the consistency score recorded by the scoring calculation module and the threshold.
[0087] The system packages the risk report along with the user-uploaded raw video data, accident collision parameters, and geographic data, and pushes it to the manual review system's work queue through a preset interface. The data is sorted by risk level and submission time for priority processing by reviewers. After logging into the backend system, reviewers can view complete case information and automatically marked suspicious locations. They then combine their professional experience to conduct on-site inspections and video reviews to make a final decision on whether to approve the claim, request supplementary materials, or reject the claim.
[0088] If all three judgment results indicate a non-fraudulent claim, the claim is marked as non-fraudulent and automatically processed. Specifically, a non-fraudulent claim refers to a claim that, after cross-verification using a voiceprint matching model, image matching model, and scoring calculation model, does not trigger a fraud warning and has a comprehensive risk level below a preset threshold. Automatic claim processing refers to an automated business process where the system automatically triggers the compensation payment process based on the case's damage assessment results, transferring the compensation funds to the user's designated account without human intervention.
[0089] In practice, after verifying the claims across three dimensions—voiceprint matching, image energy logic matching, and consistency scoring—the system summarizes the results from each module. When all three verification modules output that the claim is non-fraudulent, the system marks it as such. Subsequently, the system outputs the approved compensation amount based on a preset compensation algorithm. The system then assembles the case number, user identity information, approved compensation amount, and receiving account information into a payment request message, which is sent to the insurance core payment system via a secure interface, triggering the compensation transfer. After payment is completed, the system pushes a claim closure notification to the user's mobile terminal and updates the case status to "completed."
[0090] This embodiment only requires any judgment result indicating a fraudulent claim to mark the claim as fraudulent, thereby improving fraud detection efficiency. Furthermore, by establishing a tiered processing mechanism combining automated system screening with professional human assessment, high-risk suspicious cases receive focused attention and professional verification. This leverages the rapid filtering capabilities of multimodal models across massive numbers of cases while retaining the flexibility for human judgment in complex cases, effectively reducing the false positive rate and significantly alleviating the workload of manual review, thus lowering costs. Simultaneously, by establishing an automated claims processing channel for low-risk cases, normal claims can be processed automatically from submission to payment within seconds, significantly improving user experience and operational efficiency. Moreover, by concentrating human review resources on high-risk suspicious cases, it optimizes the allocation of review resources and reduces costs.
[0091] This application first acquires and preprocesses user-uploaded claim applications containing accident videos, vehicle and accident information, and geographical location. Then, it performs three independent verifications: First, it extracts collision sound features using a voiceprint matching model and compares them with a standard acoustic library to verify the authenticity of the collision sound. Second, it constructs a vehicle damage model using an image matching model and calculates the collision absorption energy to verify the physical logic of energy and deformation. Third, it calculates audio-visual synchronization and environmental matching scores using a scoring model, and weights these scores to obtain a cross-modal consistency score. Applications deemed fraudulent by any of these verifications are automatically transferred to a manual review queue. Applications that pass all three verifications trigger fully automated claims processing, completing the entire process of damage assessment, claims review, and payment. This method maintains near-second-level claims processing efficiency while achieving multi-dimensional fraud risk interception, significantly improving the accuracy and efficiency of fraud detection in claims, thus enhancing the accuracy, security, and reliability of insurance company claims processing, and reducing the amount of manual review to lower costs.
[0092] Figure 8 This is a schematic block diagram of a cross-modal claims fraud identification device 300 provided in an embodiment of the present invention. Figure 8 As shown, corresponding to the above-described cross-modal claims fraud identification method, the present invention also provides a cross-modal claims fraud identification device 300. The cross-modal claims fraud identification device 300 includes a unit for executing the above-described cross-modal claims fraud identification method, and this device can be configured in a computer device. Specifically, please refer to... Figure 8 The cross-modal fraud detection device 300 includes an acquisition unit 301, a voiceprint matching unit 302, an image matching unit 303, a scoring calculation unit 304, and a marking unit 305.
[0093] The acquisition unit 301 is used to acquire the claim application uploaded by the user, which includes video data of the vehicle accident, accident collision parameters and geographical data; Voiceprint matching unit 302 is used to determine whether the claim application is a fraudulent application based on the video data, the accident collision parameters and the preset acoustic library according to the preset voiceprint matching model, and output a first judgment result; The image matching unit 303 is used to determine whether the video data matches the accident collision parameters based on a preset image matching model, determine whether the claim application is a fraudulent application based on the matching result, and output a second judgment result. The scoring calculation unit 304 is used to calculate a consistency score based on the video data and the geographic data according to a preset scoring calculation model, and to determine whether the claim application is a fraudulent application based on the consistency score and a preset scoring threshold, and output a third judgment result. The marking unit 305 is used to mark the claim application as a fraudulent application if any of the first judgment result, the second judgment result, and the third judgment result indicates that the claim application is a fraudulent application.
[0094] In one embodiment, the voiceprint matching unit 302 includes a separation extraction subunit, a matching subunit, a comparison subunit, and a first determination subunit.
[0095] A separation and extraction subunit is used to separate and process the video data to generate audio data, and to extract the voiceprint features of the audio data; The matching subunit is used to match the collision acoustic features of the corresponding vehicle from the preset acoustic library according to the accident collision parameters. A comparison subunit is used to compare the voiceprint features and the collision acoustic features; The first determination subunit is used to determine that the claim application is a fraudulent application if the voiceprint feature and the collision acoustic feature do not match, and outputs the first determination result that the claim application is a fraudulent application.
[0096] In one embodiment, the separation and extraction subunit includes a filtering subunit, a conversion subunit, and an identification subunit.
[0097] A filtering subunit is used to perform multi-source noise filtering on the audio data to obtain filtered audio. A conversion subunit is used to convert the filtered audio into a frequency domain spectrum. The identification subunit is used to identify the frequency peak and decay time curve of the frequency domain spectrum.
[0098] In one embodiment, the image matching unit 303 includes a frame separation subunit, a damage modeling subunit, an energy determination subunit, and a second determination subunit.
[0099] The frame separation subunit is used to separate the video data to generate video frame data; The damage modeling subunit is used to perform damage modeling and calculate the absorbed energy during the collision based on the video frame data and the vehicle parameters. An energy determination subunit is used to calculate the theoretical kinetic energy reference value of the colliding vehicle speed and the colliding material, and to determine whether the theoretical kinetic energy reference value matches the absorbed energy. The second determination subunit is used to determine that the claim application is a non-fraudulent application if a match is found, and output the second determination result that the claim application is a non-fraudulent application; if a match is not found, the claim application is determined to be a fraudulent application, and output the second determination result that the claim application is a fraudulent application.
[0100] In one embodiment, the scoring calculation unit 304 includes an audio-visual separation subunit, an audio-visual scoring subunit, an environmental scoring subunit, and a weighted calculation subunit.
[0101] The audio-visual separation subunit is used to separate and process the video data to generate audio data and video frame data; The audio-visual scoring subunit is used to calculate the audio-visual matching score based on the audio data and the video frame data. An environmental scoring subunit is used to calculate an environmental matching score between the audio data and the geographic data. The weighted calculation subunit is used to calculate the consistency score by weighting the audio-visual matching score and the environmental matching score.
[0102] In one embodiment, the audio-visual scoring subunit includes a time point extraction subunit, a start point extraction subunit, an offset calculation subunit, and an audio-visual calculation subunit.
[0103] The time point extraction subunit is used to extract the time points of the collision sound pulses based on the audio data. The starting point extraction subunit is used to extract the starting point of vehicle displacement or shaking based on the image frame data; The offset calculation subunit is used to calculate the time offset between the time point and the starting point based on a preset timestamp alignment algorithm. The audio-visual calculation subunit is used to calculate the audio-visual matching score based on the time offset.
[0104] In one embodiment, the environmental scoring subunit includes a reverberation extraction subunit and an environmental calculation subunit.
[0105] A reverberation extraction subunit is used to extract the reverberation features of the audio data; An environmental calculation subunit is used to calculate the environmental matching score between the reverberation features and the geographic data.
[0106] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned cross-modal claims fraud identification device and its various units can be referred to the corresponding descriptions in the foregoing method embodiments. For the sake of convenience and brevity, these details will not be repeated here.
[0107] The aforementioned cross-modal fraud detection device 300 can be implemented as a computer program, which can, for example... Figure 9 It runs on the computer device shown.
[0108] Please see Figure 9 , Figure 9This is a schematic block diagram of a computer device provided in an embodiment of this application. The computer device 500 can be a terminal or a server. The terminal can be an electronic device with communication functions, such as a smartphone, tablet, laptop, desktop computer, personal digital assistant, or wearable device. The server can be a standalone server or a server cluster composed of multiple servers.
[0109] See Figure 9 The computer device 500 includes a processor 502, a memory, and a network interface 505 connected via a system bus 501. The memory may include a non-volatile storage medium 503 and internal memory 504.
[0110] The non-volatile storage medium 503 may store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions that, when executed, cause the processor 502 to perform a cross-modal claims fraud identification method.
[0111] The processor 502 provides computing and control capabilities to support the operation of the entire computer device 500.
[0112] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute a cross-modal claims fraud identification method.
[0113] This network interface 505 is used for network communication with other devices. Those skilled in the art will understand that... Figure 9 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device 500 to which the present application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0114] The processor 502 is used to run a computer program 5032 stored in the memory to implement the steps of the above-described cross-modal fraud detection method.
[0115] It should be understood that in the embodiments of this application, the processor 502 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0116] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program includes program instructions and can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.
[0117] Therefore, the present invention also provides a storage medium. This storage medium can be a computer-readable storage medium. The storage medium stores a computer program, wherein the computer program includes program instructions. When executed by a processor, the program instructions cause the processor to perform the steps of the above-described cross-modal claims fraud identification method.
[0118] The storage medium can be any computer-readable storage medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk.
[0119] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0120] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of each unit is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0121] The steps in the method of this invention can be adjusted, merged, or reduced in order according to actual needs. The units in the device of this invention can be merged, divided, or reduced according to actual needs. Furthermore, the functional units in the various embodiments of this invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0122] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0123] It should be noted that any AI models, software tools, or components not belonging to this company appearing in the embodiments of this application are merely illustrative examples and do not represent actual use. All user personal information involved in the embodiments of this application has been authorized (with the knowledge and consent) by the relevant parties or has been fully authorized by all parties, and the executing entity may obtain it through various legal and compliant means. The collection, storage, use, processing, transmission, provision, and disclosure of the information, data, and signals involved all comply with relevant laws and regulations and do not violate public order and good morals.
[0124] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for identifying claims fraud based on cross-modal methods, characterized in that, The method includes: Obtain the claim application uploaded by the user, which includes video data of the vehicle accident, accident collision parameters, and geographical data; Based on a preset voiceprint matching model, the system determines whether the claim application is a fraudulent application according to the video data, the accident collision parameters, and a preset acoustic library, and outputs a first judgment result. Based on a preset image matching model, determine whether the video data matches the accident collision parameters, determine whether the claim application is a fraudulent application based on the matching result, and output a second judgment result; Based on a preset scoring calculation model, a consistency score is calculated according to the video data and the geographic data. Based on the consistency score and a preset scoring threshold, it is determined whether the claim application is a fraudulent application, and a third judgment result is output. If any of the first, second, and third judgment results in the claim application being a fraudulent application, then the claim application will be marked as a fraudulent application.
2. The method according to claim 1, characterized in that, The step of determining whether the claim application is a fraudulent application based on the preset voiceprint matching model according to the video data, the accident collision parameters and the preset acoustic library, and outputting a first judgment result includes: The video data is separated to generate audio data, and the voiceprint features of the audio data are extracted. Based on the accident collision parameters, the collision acoustic features of the corresponding vehicle are matched from the preset acoustic library. Compare the voiceprint features with the collision acoustic features; If the voiceprint features and the collision acoustic features do not match, the claim application is determined to be a fraudulent application, and the first determination result that the claim application is a fraudulent application is output.
3. The method according to claim 2, characterized in that, The voiceprint features include frequency peaks and decay time curves. The steps of separating and processing the video data to generate audio data and extracting the voiceprint features from the audio data include: The audio data is subjected to multi-source noise filtering to obtain filtered audio; Convert the filtered audio into a frequency domain spectrum; Identify the frequency peak and decay time curve of the frequency domain spectrum.
4. The method according to claim 1, characterized in that, The accident collision parameters include vehicle parameters, collision speed, and collision material. The step of determining whether the video data matches the accident collision parameters based on a preset image matching model, determining whether the claim application is a fraudulent application based on the matching result, and outputting a second determination result includes: The video data is separated to generate image frame data; Damage modeling is performed based on the image frame data and the vehicle parameters, and the absorbed energy during the collision is calculated. Calculate the theoretical kinetic energy reference value of the colliding vehicle speed and the colliding material, and determine whether the theoretical kinetic energy reference value matches the absorbed energy; If a match is found, the claim is determined to be a non-fraudulent claim, and the second determination result that the claim is a non-fraudulent claim is output. If there is no match, the claim application is determined to be a fraudulent application, and the second determination result that the claim application is a fraudulent application is output.
5. The method according to claim 1, characterized in that, The consistency score includes an audio-visual matching score and an environmental matching score. The steps of calculating the consistency score based on the video data and the geographic data using a preset scoring calculation model include: The video data is separated to generate audio data and image frame data; The audio-visual matching score is calculated based on the audio data and the image frame data. Calculate the environmental matching score between the audio data and the geographic data; The audio-visual matching score and the environmental matching score are weighted to calculate the consistency score.
6. The method according to claim 5, characterized in that, The step of calculating the audio-visual matching score based on the audio data and the image frame data includes: Extract the time points of the collision sound pulses from the audio data; Extract the starting point of vehicle displacement or shaking based on the image frame data; The time offset between the time point and the starting point is calculated based on a preset timestamp alignment algorithm; The audio-visual matching score is calculated based on the time offset.
7. The method according to claim 5, characterized in that, The step of calculating the environmental matching score based on the audio data and the geographic data includes: Extract the reverberation features of the audio data; Calculate the environmental matching score between the reverberation features and the geographic data.
8. A cross-modal claims fraud identification device, characterized in that, Includes a unit for performing the method as described in any one of claims 1-7.
9. A computer device, characterized in that, The computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method as described in any one of claims 1-7.
10. A storage medium, characterized in that, The storage medium stores a computer program, which includes program instructions that, when executed by a processor, can implement the method as described in any one of claims 1-7.