A Method and System for Tensor Feature Extraction and Early Warning of Multimodal Data in Movable Property Pledge

CN122571518APending Publication Date: 2026-08-14FUDAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-20
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0006]为了解决现有技术存在由于多模态特征融合采用侧重共识的顺向提取逻辑而缺乏对跨模态矛盾特征的深度挖掘机制,导致局部篡改异常被正常模态数据平滑掩盖且难以捕捉具有强欺诈指示性的细粒度矛盾特征,进而对蓄意伪造数据的深层欺诈风险识别与预警能力不足的技术问题,本发明实施例提供了动产质押多模态数据张量特征提取及预警方法及其系统

Benefits of technology

[0009] The beneficial effects of the technical solutions provided by the embodiments of the present invention include at least the following:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122571518A_ABST
    Figure CN122571518A_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for tensor feature extraction and early warning of multimodal data in movable asset pledge, belonging to the field of multimodal data processing technology. The method first acquires multimodal data to construct a high-order tensor; based on this, a consensus reconstruction tensor is obtained; next, the difference between the original and reconstructed tensors is calculated to form a global residual tensor, and residual sub-tensors are extracted according to modality slices; further, a cross-modal interactive covariance tensor is constructed, and eigenvalue distribution spectrum dispersion index is extracted by feature decomposition as a cross-modal contradiction feature; finally, the consensus core tensor and contradiction features are introduced into a risk scoring network, with the consensus core mapping the benchmark risk, and the contradiction features serving as nonlinear penalty factors to superimpose a fraud risk score and trigger an early warning. This invention accurately captures fine-grained contradictions between modalities through residual decoupling and cross-covariance, avoiding the smoothing of single-point tampering anomalies, and significantly improving the ability to identify and warn against adversarial deep fraud such as forged data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multimodal data processing technology, and in particular to a method and system for extracting tensor features and providing early warning for multimodal data on pledged movable property. Background Technology

[0002] Movable asset pledging, as an important means of corporate financing, is characterized by high liquidity, large value fluctuations, and high regulatory difficulty, making its risk management a key focus for financial institutions. Meanwhile, the widespread application of the Internet of Things, warehouse management systems, logistics tracking systems, and enterprise information systems has generated a large amount of multimodal data from diverse sources in movable asset pledging transactions. Therefore, how to comprehensively, accurately, and promptly identify and warn of risks in movable asset pledging transactions has become a crucial issue that urgently needs to be addressed in the field of financial risk control.

[0003] Existing technologies typically begin by collecting multimodal data, including enterprise information, financial data, inventory and logistics records, contract texts, and surveillance images. This data is then transformed into a unified digital representation through missing value processing, normalization, text vectorization, and image / video feature extraction. Subsequently, data from different modalities and time dimensions are constructed into high-order tensors to preserve the correlation features between multi-source information. Potential risk features are then extracted using tensor decomposition or deep learning models. After fusion analysis, the extracted multimodal features are input into machine learning or deep neural network models for risk assessment, calculating risk scores and classifying risk levels. When the score exceeds a preset threshold, an early warning is triggered, simultaneously generating risk sources and disposal suggestions. This enables dynamic monitoring and early intervention of movable asset pledge risks, thus forming a complete technical process from data collection to early warning output.

[0004] However, in the process of implementing the inventive technical solution in the embodiments of this application, it was found that the above-mentioned technology has at least the following technical problems:

[0005] Movable asset pledging is susceptible to adversarial risks such as double pledging and short-selling pledging. Single-modal data is easily tampered with (e.g., falsifying IoT sensing data to conceal the true inventory status), leading to deep semantic inconsistencies and feature exclusions between multimodal data (e.g., IoT sensing modality and visual reality modality). However, existing multimodal feature fusion technologies mainly adopt a "forward" extraction logic, focusing on extracting consensus features between multi-source data for risk assessment. Under this mechanism, anomalies in local modalities are often smoothed or masked by normal data from other modalities, lacking a deep mining and detection mechanism for contradictory features between cross-modal data. This fusion paradigm, which focuses on data consistency, has limitations when facing adversarial scenarios involving data tampering. It struggles to effectively capture fine-grained contradictory features with strong fraud indicative power, such as "mismatch between IoT data and visual data," from tensor feature extraction, resulting in insufficient ability to identify and warn of deep-seated fraud risks caused by intentional falsification of single-point data. Summary of the Invention

[0006] To address the technical problems of existing technologies, which lack in-depth mining mechanisms for cross-modal contradictory features due to the consensus-oriented forward extraction logic used in multimodal feature fusion, resulting in local tampering anomalies being smoothly masked by normal modal data and difficulty in capturing fine-grained contradictory features with strong fraud indices, thus hindering the identification and early warning capabilities for deep-seated fraud risks of deliberately forged data, this invention provides a method and system for tensor feature extraction and early warning of movable property pledge multimodal data. The technical solution is as follows:

[0007] On the one hand, a method for tensor feature extraction and early warning of multimodal data in movable asset pledge is provided. This method includes: acquiring multimodal data from movable asset pledge business, where the multimodal data includes at least IoT sensing modal data and visual real-time modal data; performing feature mapping and temporal dimension alignment on the multimodal data to construct a multimodal high-order tensor containing both temporal and modal dimensions; performing low-rank tensor decomposition on the multimodal high-order tensor to extract a consensus core tensor reflecting cross-modal correlation features between multi-source data and factor matrices corresponding to each modality, and performing tensor multiplication operations along each modal dimension on the consensus core tensor and factor matrices to reconstruct the consensus reconstruction tensor; calculating the difference between the multimodal high-order tensor and the consensus reconstruction tensor to obtain a global residual tensor, and extracting features from the global residual tensor by slicing according to the modal dimension. The system employs an IoT sensing residual sub-tensor and a visual reality residual sub-tensor. It calculates the cross-covariance between these two tensors to construct a cross-modal interaction covariance tensor. The cross-modal interaction covariance tensor undergoes eigenvalue decomposition to extract the eigenvalue distribution spectrum. The system then calculates the dispersion index between the principal eigenvalue and other eigenvalues ​​in the eigenvalue distribution spectrum, identifying this dispersion index as the cross-modal contradiction feature. This cross-modal contradiction feature, along with the consensus core tensor, is introduced into a risk scoring network to output a fraud risk score. When the fraud risk score exceeds a preset threshold, a movable property pledge risk warning is triggered. The risk scoring network maps the consensus core tensor to a baseline risk value and uses the cross-modal contradiction feature as a penalty adjustment factor to perform nonlinear penalty superposition on the baseline risk value, resulting in the fraud risk score.

[0008] On the other hand, a system for extracting and warning of tensor features from movable asset pledge multimodal data is provided, including: a multimodal tensor construction module, a tensor consensus decomposition module, a modal residual extraction module, a contradiction feature extraction module, and a pledge risk warning module. The multimodal tensor construction module is used to acquire multimodal data from movable asset pledge business, including at least IoT sensing modal data and visual real-time modal data. It performs feature mapping and time dimension alignment on the multimodal data to construct a multimodal high-order tensor containing both time and modal dimensions. The tensor consensus decomposition module performs low-rank tensor decomposition on the multimodal high-order tensor, extracting a consensus core tensor reflecting cross-modal correlation features between multi-source data and factor matrices corresponding to each modality. It then performs tensor multiplication operations along each modal dimension on the consensus core tensor and factor matrices to reconstruct the consensus reconstruction tensor. The modal residual extraction module calculates the difference between the multimodal high-order tensor and the consensus reconstruction tensor. The system extracts the IoT sensing residual sub-tensor and the visual reality residual sub-tensor from the global residual tensor according to the modal dimension. A contradiction feature extraction module calculates the cross-covariance between the IoT sensing residual sub-tensor and the visual reality residual sub-tensor to construct a cross-modal interaction covariance tensor. The cross-modal interaction covariance tensor undergoes feature decomposition to extract the eigenvalue distribution spectrum and calculates the dispersion index between the principal eigenvalue and other eigenvalues ​​in the eigenvalue distribution spectrum, identifying the dispersion index as the cross-modal contradiction feature. A pledge risk warning module introduces the cross-modal contradiction feature and the consensus core tensor into a risk scoring network to output a fraud risk score. When the fraud risk score exceeds a preset threshold, a movable asset pledge risk warning is triggered. The risk scoring network maps the consensus core tensor to a benchmark risk value and uses the cross-modal contradiction feature as a penalty adjustment factor to perform nonlinear penalty superposition on the benchmark risk value to obtain the fraud risk score.

[0009] The beneficial effects of the technical solutions provided by the embodiments of the present invention include at least the following:

[0010] This invention addresses adversarial risk scenarios in movable asset pledge transactions, such as duplicate pledging, empty pledges, and single-point data forgery. It breaks through the traditional multimodal fusion method's technical approach, which only focuses on the consistency and consensus features of multi-source data. Instead of relying solely on the collaborative information between different modalities for risk assessment, it further explores the contradictory relationships between cross-modal data based on the extraction of consensus features. This effectively solves the problem in existing technologies where local abnormal data is smoothly masked by other normal modal data, making it difficult to detect deep-seated fraud risks.

[0011] Specifically, this invention first performs unified feature mapping and temporal alignment on multimodal data such as IoT sensing data and visual real-time data, and constructs a multimodal high-order tensor to realize the correlation representation of heterogeneous data in a unified space. Compared with the existing technology that processes different modalities separately or uses simple feature splicing, this invention can more completely preserve the temporal correlation and structural features between multi-source data, providing a reliable data foundation for subsequent deep feature analysis.

[0012] Building upon this foundation, this invention extracts the consensus core tensor and modality factor matrices through low-rank tensor decomposition, and utilizes a reconstruction mechanism to obtain consistency features in multimodal data. Unlike existing technologies that directly use fused features for risk assessment, this invention further utilizes the difference information between the original tensor and the consensus reconstruction tensor to construct a global residual tensor, thereby preserving anomalous information masked by consensus features. This allows potential risk features to be effectively separated from background noise, improving the identifiability of anomalous features.

[0013] Furthermore, this invention constructs a cross-modal interactive covariance tensor by performing cross-covariance analysis on the IoT sensing residual subtensor and the visual reality residual subtensor, and extracts cross-modal contradictory features based on the eigenvalue distribution spectrum. Compared to existing technologies that mainly focus on the degree of consistency between modalities, this invention can proactively identify semantic conflicts, state mismatches, and feature exclusivity phenomena between different modalities. For example, when IoT sensing data shows a normal inventory status while visual reality data reflects an abnormal inventory, this invention can effectively capture such abnormal correlations through cross-modal contradictory features, thereby discovering fine-grained risk signals with strong fraud indicativeness and improving the ability to identify intentional data forgery, partial data tampering, and concealed fraudulent activities.

[0014] Furthermore, this invention incorporates a consensus core tensor characterizing the overall business status and cross-modal conflict features reflecting the degree of abnormal conflict into the risk scoring network. By using cross-modal conflict features as a nonlinear penalty adjustment factor, the baseline risk value is dynamically corrected. Compared to existing technologies that rely solely on fused features for risk assessment, this invention simultaneously considers both the overall business status and the degree of cross-modal abnormal conflict. This allows the risk scoring results to not only reflect normal business fluctuations but also demonstrate greater sensitivity to potential fraudulent activities, thereby improving the accuracy and reliability of risk assessment.

[0015] Therefore, by constructing a technical chain of "consensus feature extraction - residual anomaly retention - cross-modal contradiction mining - risk penalty enhancement", this invention extends from multimodal data consistency analysis to multimodal data contradiction analysis. It can effectively overcome the shortcomings of existing technologies in utilizing cross-modal contradiction information, improve the ability to identify and warn of deep fraud risks in adversarial tampering scenarios, and enhance the level of intelligence in movable property pledge business risk management. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 A flowchart illustrating the method for extracting tensor features and providing early warning for multimodal data of movable property pledge provided in this application embodiment;

[0018] Figure 2 A diagram illustrating the overall logic of the method for extracting tensor features from multimodal data of movable property pledge provided in this application embodiment;

[0019] Figure 3 This is a schematic diagram of the structure of the movable property pledge multimodal data tensor feature extraction and early warning system provided in the embodiments of this application. Detailed Implementation

[0020] Embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of the present disclosure are shown in the drawings, it should be understood that embodiments of the present disclosure may be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure.

[0021] It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure. In the description of the embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "this embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc., may refer to different or the same objects.

[0022] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0023] With the rapid development of supply chain finance and warehousing supervision, movable asset pledging has become an important credit enhancement method for corporate financing. Financial institutions typically monitor the quantity, condition, and storage status of pledged goods to reduce credit risk and ensure the safety of pledged assets. To improve supervision efficiency and reduce manual verification costs, the industry is gradually introducing digital technologies such as IoT sensing devices, video surveillance equipment, and image recognition systems to collect and analyze the inventory status, location changes, and on-site conditions of pledged goods in real time, forming a multimodal business data system that includes IoT sensing data and visual real-time data. Integrating multi-source data for risk monitoring and early warning has become an important development direction for intelligent supervision of movable asset pledging.

[0024] However, in real-world business scenarios, movable asset pledging still faces risks such as double pledging, empty pledges, fictitious warehousing, and deliberate tampering with regulatory data. Particularly in adversarial fraud scenarios, malicious actors may forge or tamper with data from a single source. For example, they might use falsified inventory data uploaded by IoT devices to conceal real inventory changes, or use historical images to replace real-time monitoring footage to create the illusion of normal operations. Because different modalities of data originate from different collection channels, their authenticity and reliability vary. When some modalities of data are maliciously manipulated, it often leads to hidden semantic conflicts and inconsistencies between IoT sensing data and visual real-time data. These cross-modal contradictions usually contain strong risk indicators and are important evidence for identifying deep-seated fraudulent activities.

[0025] Existing technologies in multimodal risk analysis typically employ feature splicing, weighted fusion, attention fusion, or deep feature fusion to aggregate consistency information from different modalities and conduct risk assessments based on the fusion results. These methods primarily focus on uncovering collaborative relationships and consensus features among multi-source data to enhance overall feature representation capabilities. However, when a modality of data undergoes local tampering or anomalies, the abnormal information is easily diluted or smoothed by a large amount of normal data during the fusion process because other modalities remain normal. This makes it difficult to effectively retain anomalous features with significant risk value. Furthermore, existing methods generally lack specialized modeling mechanisms for cross-modal conflict relationships, making it difficult to deeply analyze the degree of conflict and its evolution between different modalities. Consequently, it is difficult to discover fine-grained risk features with strong fraud indicative power, such as "normal IoT sensing data but abnormal visual data" or "normal visual data but abnormal inventory sensing data."

[0026] Therefore, how to not only extract consensus features reflecting the overall business status during multimodal data fusion, but also effectively retain and mine contradictory information between different modalities, accurately identify deep anomalies caused by local tampering, data forgery, etc., and improve the ability to identify and warn of adversarial risks such as double pledging and empty pledges, has become a critical technical problem that urgently needs to be solved in the field of intelligent risk management of movable asset pledges. Based on this, proposing a method for tensor feature extraction and early warning of multimodal data in movable asset pledges that can simultaneously take into account multimodal consensus feature extraction and cross-modal contradictory feature mining has significant practical significance and application value.

[0027] like Figure 1 The diagram shown is a flowchart of the method for extracting and providing early warning of multimodal data tensors for movable property pledge provided in this application embodiment. Figure 1 As can be seen, this invention acquires multimodal data, including IoT sensing modality and visual real-time modality, from movable property pledge transactions, performs feature mapping and time alignment, and constructs a high-order tensor containing time and modality dimensions to achieve unified representation and collaborative analysis of heterogeneous data. Subsequently, it extracts the consensus core tensor and modality factor matrices through low-rank tensor decomposition and performs tensor reconstruction to effectively extract cross-modal correlation features of multimodal data and eliminate noise and redundant information. Based on this, it calculates the residual between the original tensor and the consensus-reconstructed tensor and slices them by modality to retain local anomalies masked by consensus fusion. This invention identifies modal deviations to provide a basis for locating potential fraud. Furthermore, it constructs a cross-modal interactive covariance tensor and extracts the dispersion of principal eigenvalues ​​through feature decomposition, forming cross-modal contradictory features. This allows for in-depth mining of semantic conflicts and anomalous signals between different modalities, improving sensitivity to local tampering and data forgery. Finally, the consensus core tensor and cross-modal contradictory features are input into a risk scoring network. A nonlinear penalty mechanism generates a fraud risk score, triggering a risk warning when the score exceeds a threshold. This enables accurate identification and early warning of adversarial risks such as duplicate pledging, empty pledging, and forged data. Through these technical means, this invention not only effectively utilizes cross-modal correlation features between multimodal data but also highlights and quantifies contradictory features between modalities, achieving high-precision identification and timely warning of adversarial fraud behaviors such as duplicate pledging, empty pledging, and data tampering. This significantly improves the accuracy and intelligence level of movable asset pledge risk management.

[0028] refer to Figure 1As the first step in the method for tensor feature extraction and early warning of multimodal data in movable asset pledge, the specific steps are as follows: First, acquire multimodal data of movable asset pledge business, including at least IoT sensing modal data and visual real-time modal data. Second, perform feature mapping and temporal dimension alignment on the multimodal data to construct a high-order multimodal tensor containing both temporal and modal dimensions. This step maps heterogeneous data from different sources and with different structures to a unified tensor space, achieving synchronous association and unified expression of multimodal data in the temporal dimension. This provides a standardized data foundation for subsequent cross-modal collaborative analysis, avoids information distortion caused by differences in data format and temporal misalignment, and improves the accuracy and completeness of data fusion.

[0029] In this embodiment, multimodal data of movable property pledge transactions are acquired, and feature mapping and time dimension alignment are performed on the multimodal data to construct a multimodal high-order tensor containing time and modal dimensions. Specifically, the steps include:

[0030] S11: Acquire heterogeneous multimodal data and perform feature mapping: Acquire multimodal data within the movable property pledge storage area, including at least IoT sensing modal data and visual real-time modal data. Specifically, IoT sensing modal data includes weight sensing time series collected by electronic weighbridges and location tracking time series collected by radio frequency identification (RFID) tags; visual real-time modal data includes image sequences of cargo space occupancy status and video frame sequences of cargo outlines collected by security cameras.

[0031] Extracting structured features from visual live modal data through feature mapping:

[0032] (1) For the image sequence of the occupancy status of the cargo space, the Otsu method or a fixed binarization threshold (such as gray value 127) is used to segment the foreground of the image. The ratio of the foreground pixels to the total pixels of the set cargo space area is calculated and mapped to the cargo space occupancy rate feature time series. The value range is [0,1].

[0033] (2) For the cargo outline video frame sequence, the Canny edge detection algorithm is used to extract the cargo outline, and the pixel area of ​​the outline-enclosed region is calculated and mapped to the outline area feature time series. .

[0034] IoT sensing modal data does not require complex mapping; the weight sensing data can be directly normalized and used as a weight feature time series. Mapping the tag coordinates read by RFID to a time series of location features .

[0035] S12: Extract initial time series and construct feature vectors: Extract the initial time series of IoT sensing modal data and visual real-time modal data respectively. For the IoT sensing modal, the same sampling time... Weight characteristics with location features Combining and constructing IoT sensing feature vectors For the visual live mode, the same sampling time will be used. Occupancy characteristics Area characteristics Combining to construct visual reality feature vectors This results in two initial time series sets: IoT sensing sequences. Visual live sequence ,in and These represent the total number of sampling points for the two sets of equipment.

[0036] S13: Calculating the optimal matching path based on a constrained dynamic time warping algorithm: Using a preset business inspection cycle (e.g., 24 hours) as the baseline time window, a multivariate dynamic time warping algorithm is used to calculate the optimal matching path between two sets of feature vector sequences. Specifically:

[0037] Define the local distance function as the Euclidean distance between two feature vectors: Construct the cumulative distance cost matrix Furthermore, Sakoe-Chiba constraints are introduced to limit the search range, and a global bandwidth threshold is set. (Preferred, This means that the offset of the alignment path cannot exceed 10% of the total length to avoid ill-conditioned matching at physically unrelated time points; dynamic programming is used to solve for the minimum cumulative cost path that satisfies the constraints. ,in The total step size after alignment. Indicates the first The sampling time index of the initial time series of the IoT sensing mode corresponding to each matching point. Indicates the first The sampling time index of the initial time series of the visual real-world modality corresponding to each matching point. For the matching point number, .

[0038] S14: Resampling and Construction of Multimodal Higher-Order Tensors:

[0039] Along the optimal matching path The IoT sensing feature vector sequence and the visual real-time feature vector sequence are resampled and aligned for fusion. For one-to-many or many-to-one mapping relationships in the path, a linear interpolation algorithm is used to calculate the feature vector values ​​at missing moments, mapping the two originally asynchronous sequences onto a unified time axis, resulting in an aligned sequence with a length of [length value missing]. The time series.

[0040] Finally, a third-order multimodal high-order tensor is constructed. :

[0041] The first dimension is time, with a size of [missing information]. The second dimension is the modality dimension, with a size of 2. The first slice is the IoT sensing modality, and the second slice is the visual real-time modality. The third dimension is the feature dimension, with a size of 2. The two-dimensional feature vectors under each modality are arranged in order (the IoT modality is [W,L], and the visual modality is [O,A]).

[0042] It should be understood that the selected IoT data (weight + location) and visual data (occupancy rate + outline) comprehensively characterize the state of the pledged items from three dimensions: physical quality, spatial location, and appearance. This fine-grained feature mapping makes it easy for forgery in a single dimension to conflict with normal features in other dimensions, providing rich semantic support for the subsequent extraction of cross-modal contradictory features and significantly improving robustness against deliberate tampering.

[0043] The second step in the method for tensor feature extraction and early warning of multimodal data in movable asset pledge is as follows: low-rank tensor decomposition is performed on the high-order tensor of the multimodality to extract the consensus core tensor reflecting the cross-modal correlation features between multi-source data and the factor matrix corresponding to each modality. The consensus core tensor and the factor matrix are then multiplied along each modal dimension to reconstruct the consensus reconstruction tensor. This step extracts the stable collaborative correlation information between different modalities through low-rank decomposition, which can effectively remove random noise and redundant interference, retain the core features reflecting the real business status, establish a reliable reference benchmark for subsequent anomaly detection, and improve the overall feature expression capability and data utilization efficiency.

[0044] In this embodiment, low-rank tensor decomposition is performed on the multimodal high-order tensor to extract the consensus core tensor reflecting the cross-modal correlation features among multi-source data and the factor matrix corresponding to each modality. The consensus core tensor and the factor matrix are then multiplied along each modal dimension to reconstruct the consensus reconstruction tensor. The specific steps include:

[0045] S21: Modal expansion of multimodal higher-order tensors:

[0046] The High-Order Singular Value Decomposition (HOSVD) algorithm is used to transform the third-order multimodal high-order tensor constructed in the aforementioned steps. Perform a mode-n matrix expansion along each modal dimension. The expansion rule is: use the dimension of the nth modality as the row of the expanded matrix, and use the product of all remaining modal dimensions in ascending order as the column of the expanded matrix. Specifically:

[0047] Expanding along the time dimension (pattern-1) yields the time expansion matrix. The columns are composed of combinations of modal and feature dimensions; expanding along the modal dimension (mode-2) yields the modal expansion matrix. Expanding along the feature dimension (pattern-3) yields the feature expansion matrix. .

[0048] S22: Constructing modal factor matrices based on economical SVD and cumulative energy threshold: For each expanded matrix Economic singular value decomposition is performed separately to accommodate cases where the modal expansion matrix has a significantly skewed row and column ratio. ,in Left singular matrix ( For the first (dimensions of modality) It is a diagonal matrix of singular values ​​arranged in descending order.

[0049] Get the preset number corresponding to the dimensionality reduction truncated rank Identify the principal singular values, extract the associated left singular vectors, and construct the factor matrix corresponding to each mode. Among them, truncated rank The specific mathematical relationship, determined by the cumulative energy percentage method, is as follows: ;in, For the first A singular value, A preset energy retention threshold is set. Preferably, a threshold is set. .

[0050] Due to the modal dimension in this invention Feature Dimension Therefore, the cutoff rank of the modality factor matrix truncated rank of the eigenfactor matrix ; Time dimension truncated rank This yields the time factor matrix. Modal factor matrix and eigenfactor matrix .

[0051] S23: Compute the consensus core tensor: convert the multimodal high-order tensor The transpose of the factor matrix corresponding to each mode is used to perform tensor product operations (i.e., n-mode product) along the corresponding mode dimension to obtain the consensus core tensor. The specific calculation formula is as follows: ;in, Indicates along the first Modal tensor product. For ease of engineering implementation, the equivalent matrix operation for tensor product is: perform a modal-n expansion of the tensor, transpose the left factor matrix, and then fold it back into a tensor. The resulting consensus core tensor It retains only the core principal components that reflect the cross-modal correlation characteristics between multi-source data.

[0052] S24: Reconstructing computational consensus and reconstructing tensors:

[0053] Consensus Core Tensor The factor matrices corresponding to each mode are multiplied along the corresponding mode dimensions using tensor multiplication to reconstruct the consensus reconstruction tensor. The specific calculation formula is as follows: Among them, consensus reconstruction tensor It is completely consistent with the original high-order tensor dimension. Its physical meaning is: "ideal normal state" data derived from the low-rank consensus subspace of multi-source data, in which each modality corroborates each other and there is no conflict.

[0054] It should be added that the consensus reconstruction tensor obtained from the reconstruction... It is the projection of the original tensor onto the low-rank consistent subspace, representing the "ideal state" when there is no semantic conflict between IoT and visual data. This reconstruction operation makes the implicit consensus rules explicit, providing a precise mathematical benchmark for calculating the difference (i.e., residual) between the original tensor and the reconstructed tensor. It is a prerequisite for accurately identifying local tampering anomalies.

[0055] The third step in the method for extracting tensor features and providing early warning for multimodal data in movable asset pledge is as follows: Calculate the difference between the multimodal high-order tensor and the consensus reconstruction tensor to obtain the global residual tensor, and extract the IoT perception residual sub-tensor and the visual reality residual sub-tensor from the global residual tensor according to the modal dimension. This step effectively separates abnormal information from consensus information by extracting the parts of the original data that cannot be explained by the consensus model, so as to preserve the local abnormal features that are masked by the fusion process, thereby enhancing the ability to perceive inventory anomalies, data tampering and abnormal business behavior.

[0056] In this embodiment, the difference between the multimodal higher-order tensor and the consensus reconstruction tensor is calculated to obtain the global residual tensor. Then, the IoT sensing residual sub-tensor and the visual reality residual sub-tensor are extracted from the global residual tensor by slicing according to the modal dimension. The specific steps include:

[0057] S31: Calculate the global residual tensor and perform soft thresholding noise reduction: Calculate the original multimodal high-order tensor constructed in the previous steps. The consensus reconstruction tensor obtained from the reconstruction computation By subtracting element by element, we obtain the initial residual tensor: Among them, elements The positive or negative value indicates the direction (higher or lower) of the actual data deviation from the consensus benchmark. To eliminate the small fluctuations caused by inherent sensor noise, a soft thresholding function is introduced to denoise the initial residual tensor, resulting in the global residual tensor. : ;in, This is a preset residual filtering threshold. Preferably, The value is set to the mean of the absolute values ​​of the residuals of the corresponding modal features under normal, fraud-free historical data plus three times the standard deviation (i.e., (Principle); when the absolute value of the residual is less than When set to 0, only significant outliers are retained.

[0058] S32: Determine the modal dimension index mapping: in the global residual tensor In the three-order structure (time dimension × modality dimension × feature dimension), the second order is explicitly defined as the modality dimension. Based on the aforementioned feature mapping stage settings, the index mapping relationship of the modality dimension is determined as follows: index m=1 corresponds to the IoT sensing modality, and index m=2 corresponds to the visual real-time modality.

[0059] S33: Extracting Residual Subtensors by Modal Dimension Slicing: Based on the dimension index, perform dimensionality reduction slicing on the global residual tensor along the modal dimension. Fixing the index values ​​of the modal dimensions, extract all corresponding elements in the time dimension (1st order, size K) and the feature dimension (3rd order, size 2), decoupling the third-order tensor into a second-order matrix (subtensor):

[0060] Extract the coupled sensing residual tensor: fix m=1, extract The dimension is obtained as The subtensor has two columns corresponding to the weight feature residual and the position feature residual, respectively; extract the visual reality residual subtensor: fix m=2, extract The dimension is obtained as The subtensor has two columns corresponding to the occupancy feature residual and the contour area feature residual, respectively.

[0061] S34: Z-score normalization of the residual subtensors: Due to the significant differences in the dimensions and numerical magnitudes between IoT features and visual features, to prevent the dominance of large numerical features in subsequent calculations, the two extracted residual subtensors are Z-score normalized column-wise (feature dimensions). The calculation formula is as follows: ;in, For the first Column features in the first The residual values ​​at each time point and These represent the mean and standard deviation of the feature column over the time dimension, respectively. The standardized IoT sensing residual subtensor. With visual reality residual tensor All numerical values ​​follow a standard normal distribution scale, eliminating the influence of dimensions.

[0062] A typical characteristic of movable property pledge fraud is single-point forgery. If the global residuals are directly evaluated, the huge abnormal residuals of the IoT modality are easily averaged and diluted by the normal small residuals of the visual modality in the overall tensor calculation. This invention, through dimensionality reduction slicing of the modal dimension, physically severs the data compensation channel between modalities, ensuring that the tampered residuals of local single modalities are completely and independently preserved, effectively eliminating the hidden danger of abnormal signals being smoothed out by normal modalities.

[0063] The fourth step in the method for extracting and warning features of multimodal data tensors for movable asset pledge is as follows: The cross-covariance between the IoT sensing residual sub-tensor and the visual reality residual sub-tensor is calculated to construct a cross-modal interactive covariance tensor. Feature decomposition is performed on the cross-modal interactive covariance tensor to extract the eigenvalue distribution spectrum. The dispersion index between the principal eigenvalue and other eigenvalues ​​in the eigenvalue distribution spectrum is calculated, and the dispersion index is identified as the cross-modal contradictory feature. This step can identify semantic conflicts and state inconsistencies between IoT sensing results and visual reality results based on the correlation between abnormal information from different modalities. It can deeply explore hidden risks that are difficult to detect with a single modality, and improve the ability to identify adversarial fraudulent behaviors such as forged data, false inventory, and duplicate pledging.

[0064] In this embodiment, the cross-covariance between the IoT sensing residual sub-tensor and the visual reality residual sub-tensor is calculated to construct a cross-modal interaction covariance tensor. Feature decomposition is performed on the cross-modal interaction covariance tensor to extract the eigenvalue distribution spectrum. The dispersion index between the principal eigenvalue and other eigenvalues ​​in the eigenvalue distribution spectrum is calculated, and the dispersion index is determined as the cross-modal contradictory feature. Specifically, the steps include:

[0065] S41: Matrix transpose expansion and mean centering of residual subtensors: Extracting and standardizing the IoT sensing residual subtensors from the previous steps With visual reality residual tensor Expanding along the time dimension. Specifically, the subtensor is transposed, transforming the original time row and feature column structure into a matrix with features as rows and time sampling points as columns, resulting in the IoT residual matrix. With visual residual matrix .

[0066] Subsequently, mean centering is performed on both matrices. Specifically, the mean centering of each feature is calculated along the time dimension (i.e., each row). The arithmetic mean of the sample points, and the mean is subtracted from all elements in the row: ; ;in, This represents the feature index. Centering eliminates the steady-state DC bias in the residuals, ensuring that subsequent covariance reflects only the dynamic fluctuations between features.

[0067] S42: Constructing the cross-modal interaction covariance matrix and tensor reshaping: Calculate the matrix product between the mean-centered IoT residual matrix and the transpose of the visual residual matrix to obtain the cross-modal interaction covariance matrix. : ;in, Its elements Represents the first IoT sensing The feature residuals and visual reality The cross-covariance matrix is ​​calculated based on the cross-modal interaction covariance. By reshaping along the feature dimension, we obtain the second-order cross-modal interaction covariance tensor.

[0068] S43: Extracting the eigenvalue distribution spectrum based on singular value decomposition:

[0069] Due to the cross-modal interaction covariance matrix Since the matrix is ​​asymmetric, direct eigenvalue decomposition will produce complex eigenvalues, making sorting and comparison impossible. Therefore, the Singular Value Decomposition (SVD) algorithm is used to decompose the matrix corresponding to the covariance tensor. Decompose: Extract the diagonal matrix The singular values ​​on the spectrum are used as the distribution spectrum of generalized eigenvalues. Because... After decomposition, two non-negative real singular values ​​are obtained. These are then arranged in descending order of their numerical values ​​to obtain the eigenvalue distribution spectrum. .

[0070] S44: Calculate the dispersion index and identify it as a cross-modal contradictory feature:

[0071] Extract the principal eigenvalue with the largest value from the eigenvalue distribution spectrum. Calculate the mean of all eigenvalues ​​in the eigenvalue distribution spectrum except for the principal eigenvalues ​​to obtain the background energy of the secondary features. (In this embodiment, That is The ratio between the principal eigenvalue and the secondary eigenvalue background energy is calculated as the spectral gap parameter. To prevent computational overflow caused by minor eigenvalues ​​tending to zero under normal conditions, a minimal constant is introduced. (Preferred, Perform denominator smoothing: ; Spectral gap parameters The dispersion index, used to characterize the degree of local feature mutation caused by data tampering, is identified as a cross-modal contradictory feature. In practical business judgment, the spectral gap parameter is used... Contradiction determination threshold Comparison; Contradiction Detection Threshold Based on the statistical standard of historical data on fraud-free movable property pledges, preferably, the following is set: ,when When this occurs, it indicates that the energy of the principal eigenvalues ​​far exceeds the background energy, and there is a serious one-sided deviation and contradiction between modes.

[0072] A cross-modal interaction covariance tensor was constructed by calculating the transpose product of the IoT residual and the visual residual. This operation breaks through the limitations of single-modal internal evaluation, directly comparing whether the "abnormal fluctuations in IoT perception" and the "abnormal fluctuations in visual reality" are synchronized at the feature level. When the data of a certain modality is deliberately tampered with, its residual changes will manifest as a serious mismatch and deviation from the residuals of other modalities in the covariance matrix, thus transforming the hidden single-point tampering into an explicit cross-modal conflict.

[0073] The fifth step in the method for extracting tensor features and providing early warnings for multimodal data in movable asset pledge is as follows: Cross-modal conflict features and the consensus core tensor are introduced into a risk scoring network to output a fraud risk score. When the fraud risk score exceeds a preset threshold, a movable asset pledge risk warning is triggered. Specifically, the risk scoring network maps the consensus core tensor to a baseline risk value and uses cross-modal conflict features as a penalty adjustment factor to perform nonlinear penalty superposition on the baseline risk value to obtain the fraud risk score. This step integrates overall business status information and cross-modal abnormal conflict information into the risk assessment process, enabling dynamic quantitative analysis of risk levels. This allows the risk score to not only reflect the business operation status but also sensitively respond to changes in abnormal conflicts, thereby improving the accuracy of risk identification and the timeliness of early warnings, and enhancing the security management capabilities of movable asset pledge business.

[0074] In this embodiment, cross-modal contradiction features and consensus core tensors are introduced into the risk scoring network to output a fraud risk score. When the fraud risk score exceeds a preset threshold, a movable property pledge risk warning is triggered, specifically including the following steps:

[0075] S51: Global Temporal Pooling and Vectorization Preprocessing of the Consensus Core Tensor: Due to the consensus core tensor extracted in the previous steps... In the middle, the time dimension truncates the rank Since the input time window length dynamically changes, it cannot be directly input into a neural network with a fixed topology. Therefore, it is first processed along the time dimension (the first dimension). Perform global average pooling. Its physical meaning is to extract the steady-state consensus features over the entire monitoring period, eliminate dependencies at specific time steps, and obtain the pooled low-dimensional core tensor. It was then flattened and reshaped into a one-dimensional consensus feature vector. .because , The maximum dimension of this vector is 4.

[0076] S52: Pre-training and mapping of the baseline feature extraction layer:

[0077] One-dimensional consensus feature vector The baseline feature extraction layer of the input risk scoring network is used. This layer is a multilayer perceptron (MLP), consisting of an input layer, hidden layers, and an output layer. The specific network topology is as follows: the input layer dimension is... The hidden layer dimension is set to 16 (using overcomplete mapping to enhance the feature representation capability of low-dimensional inputs), and the activation function is ReLU; the output layer dimension is 1, and the activation function is Softplus (ensuring that the output is always positive), outputting a baseline risk value in scalar form. .

[0078] It is important to note that the weights and bias parameters of this multilayer perceptron are pre-obtained through supervised learning: consensus feature vectors from historical normal operating cycles are used as training samples, and standard operational risk scores (such as risk values ​​calculated based on historical volatility) labeled by business experts are used as labels. Mean squared error (MSE) is used as the loss function for optimization and convergence. During inference, the parameters are frozen, and the output... It characterizes the level of routine operational risk in movable property pledge business under the absence of adversarial interference.

[0079] S53: Constructing an exponential penalty weight term with baseline offset and truncation:

[0080] Obtain the cross-modal contradictory features (i.e., spectral gap parameters) calculated in the aforementioned steps. Using cross-modal contradiction characteristics as the penalty adjustment factor, a preset penalty coefficient is introduced. ( To eliminate the normal state. Baseline value ( To mitigate the inherent amplification bias introduced by the exponential term and prevent system crashes due to overflow of the exponential term value during severe tampering, an exponential penalty weight term with offset and truncation is constructed. : ;in, For offset terms, ensure that the multimodal data are completely consistent ( When the exponential term approaches 0, the penalty weight approaches 1. A preset upper limit threshold for truncation is used to restrict floating-point exponentiation operations within a safe range. Preferably, a threshold is set. Preset penalty coefficient To control the sensitivity to sudden increases in conflict, preferably, set... .

[0081] S54: Nonlinear penalty superposition calculation of fraud risk score:

[0082] Exponential penalty weighting Compared with the benchmark risk value Perform multiplication to obtain the final fraud risk score. : When semantic mutual exclusion exists between multimodal data, leading to a sudden increase in cross-modal contradictory features ( When the penalty weighting factor increases exponentially, it drives the fraud risk score to rise exponentially.

[0083] S55: Triggering a risk warning for movable property pledge:

[0084] The calculated fraud risk score With the preset risk warning threshold Comparison. Risk warning threshold. The calibration method is as follows: On the validation set, calculate the historical samples that are clearly free of fraud. The distribution is used, and its 99th percentile is taken as the threshold to control the false alarm rate of normal business operations to within 1%. When At that time, it was determined that there was deliberate data tampering and adversarial fraud, triggering a risk warning for movable property pledge.

[0085] To address the issue of dynamic changes in tensor dimensions caused by variable monitoring time windows, a novel approach is introduced to extract steady-state consensus features using global average pooling, decoupling the network structure from the time step. Simultaneously, a supervised pre-training mechanism for the MLP is defined, ensuring that the output of the baseline risk value is not arbitrary but a quantitative assessment that truly aligns with the historical normal fluctuation patterns of movable asset pledge business, providing a precise foundation for subsequent penalty stacking.

[0086] Furthermore, this invention maps the consensus core tensor to the baseline risk of business "noise," and uses cross-modal contradictory features as an independent exponential penalty factor. This mechanism ensures that when normal cargo loss leads to a higher baseline risk, the scoring will not result in false alarms or fraud if there are no inconsistencies; however, once unilateral data tampering occurs, the exponential term instantly amplifies the hidden contradictory signals, achieving highly sensitive early warning with an extremely low false alarm rate based on a threshold calibrated at the 99th percentile.

[0087] In addition to the above, it is also necessary to add the following explanation: Figure 2This is an overall logic diagram of the method for extracting and warning of multimodal data tensors of movable property pledge provided in the embodiments of this application; based on Figure 2 It should be understood that the steps of this invention form a complete collaborative processing link by constructing multimodal high-order tensors, extracting consensus core tensors, calculating residual sub-tensors, mining cross-modal contradiction features, and performing risk scoring, thereby realizing closed-loop collaboration from multimodal data fusion, consistency analysis to conflict analysis and risk decision-making. First, feature mapping and temporal alignment of multimodal data ensure that different modalities are associated under a unified spatiotemporal scale, providing a unified data foundation for consensus feature extraction and cross-modal relationship analysis. Then, low-rank tensor decomposition and consensus reconstruction extract cross-modal association features of the overall business state, providing a reference baseline for residual analysis and effectively separating abnormal information from normal business information. By obtaining modal deviation information through residual slicing, highly sensitive input is provided for cross-modal contradiction analysis, allowing local anomalies hidden in the fused data to be preserved. Furthermore, cross-modal interaction covariance is constructed and contradiction features are extracted, enabling in-depth analysis from the existence of anomalies to their associations, revealing semantic conflicts and potential fraud signals between different modalities. Finally, the consensus core tensor and cross-modal contradiction features are jointly input into a risk scoring network, generating a fraud risk score through a nonlinear penalty mechanism, achieving a joint assessment of the overall business state and the degree of anomaly conflict. Each step is interdependent and mutually supportive, jointly enhancing the ability to express risk characteristics, identify abnormal behavior, and improve the accuracy of risk warnings. This effectively overcomes the problem that existing technologies rely solely on consensus information, which can easily mask abnormal information, thereby significantly improving the intelligence and precision of risk management in movable property pledge business.

[0088] like Figure 3 The diagram shown is a structural schematic of the movable property pledge multimodal data tensor feature extraction and early warning system provided in this application embodiment. (Refer to...) Figure 3 The system includes: a multi-modal tensor construction module, a tensor consensus decomposition module, a modal residual extraction module, a contradiction feature extraction module, and a pledge risk early warning module.

[0089] Specifically, the multimodal tensor construction module is used to acquire multimodal data of movable property pledge business. The multimodal data includes at least IoT sensing modal data and visual real-time modal data. Feature mapping and temporal dimension alignment are performed on the multimodal data to construct a multimodal high-order tensor containing temporal and modal dimensions.

[0090] The tensor consensus decomposition module is used to perform low-rank tensor decomposition on multimodal high-order tensors, extract the consensus core tensor reflecting the cross-modal correlation characteristics between multi-source data and the factor matrix corresponding to each modality, and perform tensor multiplication operation on the consensus core tensor and the factor matrix along each modal dimension to reconstruct the consensus reconstruction tensor.

[0091] The modal residual extraction module is used to calculate the difference between the multimodal high-order tensor and the consensus reconstruction tensor to obtain the global residual tensor. It then extracts the IoT perception residual sub-tensor and the visual reality residual sub-tensor from the global residual tensor according to the modal dimension.

[0092] The contradictory feature extraction module is used to calculate the cross-covariance between the IoT sensing residual sub-tensor and the visual reality residual sub-tensor, and construct the cross-modal interaction covariance tensor. The cross-modal interaction covariance tensor is decomposed to extract the eigenvalue distribution spectrum, and the dispersion index between the principal eigenvalue and other eigenvalues ​​in the eigenvalue distribution spectrum is calculated. The dispersion index is determined as the cross-modal contradictory feature.

[0093] The pledge risk warning module is used to introduce cross-modal contradiction features and consensus core tensors into the risk scoring network and output fraud risk scores. When the fraud risk score exceeds a preset threshold, a movable property pledge risk warning is triggered. The risk scoring network maps the consensus core tensor to a benchmark risk value and uses cross-modal contradiction features as a penalty adjustment factor to perform nonlinear penalty superposition on the benchmark risk value to obtain the fraud risk score.

[0094] In summary, this invention addresses the shortcomings of existing multimodal feature fusion technologies, which, due to their emphasis on consensus-based forward extraction logic, suffer from the smooth masking of local tampering anomalies, difficulty in capturing fine-grained cross-modal contradictions, and insufficient deep fraud early warning capabilities. It breaks through the limitations of traditional fusion paradigms by innovatively proposing a reverse extraction logic of "consensus stripping - residual decoupling - contradiction amplification." By reconstructing multimodal consensus as a control baseline, residual slicing cuts off the compensation masking of anomalies by normal data, leveraging cross-covariance and spectral decomposition to deeply mine semantic mutual exclusion between IoT and vision, and finally amplifying the risk of contradiction conflicts through a nonlinear penalty mechanism. Compared with existing technologies, this invention not only retains and mines local anomaly information masked by consensus fusion but also improves the sensitivity and accuracy of risk scoring through nonlinear adjustment of cross-modal contradiction features, thereby enhancing the early warning capability for adversarial risks such as duplicate pledging, short-selling pledging, and forged data. It has good promotional value and practical application prospects in movable asset pledge risk management and intelligent early warning applications.

[0095] Through the above description of the implementation methods, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the above functions can be divided into different functional modules to complete all or part of the functions described above.

[0096] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between devices or units, and may be electrical, mechanical, or other forms.

[0097] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units, located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0098] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0099] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to the solution, or all or part of the technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0100] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for extracting tensor features and providing early warning for multimodal data on pledged movable property, characterized in that, Includes the following steps: Acquire multimodal data of movable property pledge business, wherein the multimodal data includes at least IoT sensing modal data and visual real-time modal data; perform feature mapping and time dimension alignment on the multimodal data, and construct a multimodal high-order tensor containing time dimension and modal dimension; Low-rank tensor decomposition is performed on the multimodal high-order tensor to extract the consensus core tensor and the factor matrix corresponding to each modality, which reflect the cross-modal correlation characteristics between multi-source data. The consensus core tensor and the factor matrix are then multiplied along each modal dimension to reconstruct the consensus reconstruction tensor. The difference between the multimodal higher-order tensor and the consensus reconstruction tensor is calculated to obtain the global residual tensor. The IoT perception residual sub-tensor and the visual reality residual sub-tensor are extracted from the global residual tensor by modality dimension. The cross-covariance between the IoT sensing residual sub-tensor and the visual reality residual sub-tensor is calculated to construct a cross-modal interaction covariance tensor. The cross-modal interaction covariance tensor is decomposed into features to extract the feature value distribution spectrum. The dispersion index between the principal feature value and other feature values ​​in the feature value distribution spectrum is calculated, and the dispersion index is determined as the cross-modal contradiction feature. Cross-modal contradiction characteristics and consensus core tensors are introduced into the risk scoring network to output a fraud risk score; When the fraud risk score exceeds a preset threshold, a movable property pledge risk warning is triggered; wherein, the risk scoring network maps the consensus core tensor to a benchmark risk value, and uses cross-modal contradiction features as a penalty adjustment factor to perform nonlinear penalty superposition on the benchmark risk value to obtain the fraud risk score.

2. The method for extracting and providing early warning of multimodal data tensors for movable property pledge as described in claim 1, characterized in that, The IoT sensing modal data includes the weight sensing time series of electronic weighbridges in the pledged storage area and the location tracking time series of RFID tags; the visual real-time modal data includes the image sequence of the occupancy status of the storage space and the video frame sequence of the cargo outline collected by security cameras in the pledged storage area.

3. The method for extracting and providing early warning of multimodal data tensors for movable property pledge as described in claim 1, characterized in that, The process of feature mapping and temporal alignment of multimodal data to construct a high-order multimodal tensor containing both temporal and modal dimensions specifically includes: The initial time series of the IoT sensing modal data and the visual real-time modal data are extracted respectively; Using a preset business inspection cycle as the baseline time window, a dynamic time warping algorithm is used to calculate the optimal matching path between the initial time series of the IoT sensing modal data and the initial time series of the visual real-time modal data. The two initial time series are resampled and aligned along the optimal matching path to construct the multimodal high-order tensor with consistent time dimension.

4. The method for extracting and providing early warning of multimodal data tensors for movable property pledge as described in claim 1, characterized in that, The step of performing low-rank tensor decomposition on multimodal high-order tensors to extract consensus core tensors reflecting cross-modal correlation features among multi-source data and factor matrices corresponding to each modality specifically includes: The higher-order singular value decomposition algorithm is used to expand the multimodal higher-order tensor along each modal dimension to obtain the expansion matrix corresponding to each modal dimension. Singular value decomposition is performed on each of the expanded matrices to obtain a preset number of principal singular values ​​corresponding to the dimensionality reduction truncated rank, and the associated left singular vectors are extracted to construct the factor matrix corresponding to each mode. The consensus core tensor is obtained by performing tensor multiplication operations along each modal dimension on the multimodal higher-order tensor and the transpose of the factor matrix corresponding to each modality.

5. The method for extracting and providing early warning of multimodal data tensors for movable property pledge as described in claim 1, characterized in that, The extraction of IoT-sensing residual sub-tensors and visual reality residual sub-tensors from the global residual tensor by slicing according to the modality dimension specifically includes: In the modal dimensions of the global residual tensor, the dimension indices corresponding to the IoT sensing modal data and the visual real-time modal data are determined respectively. Based on the dimensional index, the global residual tensor is sliced ​​along the modal dimension to separate the IoT sensing residual sub-tensor containing only IoT sensing features and the visual reality residual sub-tensor containing only visual reality features.

6. The method for extracting and providing early warning of multimodal data tensors for movable property pledge as described in claim 1, characterized in that, The calculation of the cross-covariance between the IoT sensing residual sub-tensor and the visual reality residual sub-tensor to construct a cross-modal interaction covariance tensor specifically includes: The IoT sensing residual sub-tensor and the visual reality residual sub-tensor are expanded along the time dimension into IoT residual matrix and visual residual matrix, respectively. The IoT residual matrix and the visual residual matrix are subjected to mean centering. The matrix product between the mean-centered IoT residual matrix and the transpose of the visual residual matrix is ​​calculated to obtain the cross-modal interaction covariance matrix. The cross-modal interaction covariance matrix is ​​reshaped into the cross-modal interaction covariance tensor.

7. The method for extracting and providing early warning of multimodal data tensors for pledged movable property as described in claim 1, characterized in that, The dispersion index between the principal eigenvalue and other eigenvalues ​​in the calculated eigenvalue distribution spectrum is defined as a cross-modal contradictory feature, specifically including: All feature values ​​in the feature value distribution spectrum are sorted in descending order of numerical value, and the feature value with the largest numerical value is extracted as the principal feature value; Calculate the mean of all eigenvalues ​​in the eigenvalue distribution spectrum except for the principal eigenvalue to obtain the secondary feature background energy; The ratio between the principal eigenvalue and the background energy of the secondary eigenvalue is calculated as the spectral gap parameter. The spectral gap parameter is used as the dispersion index to characterize the degree of local feature mutation caused by data tampering, and is determined as the cross-modal contradictory feature.

8. The method for extracting and providing early warning of multimodal data tensors for movable property pledge as described in claim 1, characterized in that, The risk scoring network includes a baseline feature extraction layer; The risk scoring network maps the consensus core tensor to a baseline risk value, including: The consensus core tensor is input into the benchmark feature extraction layer. The consensus features contained in the consensus core tensor are then mapped to a lower dimension using a multilayer perceptron. The benchmark risk value is output in scalar form. The benchmark risk value represents the routine operational risk score of movable property pledge business under non-adversarial interference.

9. The method for extracting and providing early warning of multimodal data tensors for movable property pledge as described in claim 1, characterized in that, The fraud risk score is obtained by using cross-modal contradictory characteristics as a penalty adjustment factor to perform non-linear penalty superposition on the benchmark risk value, specifically including: Using the natural constant as the base and the product of the cross-modal contradictory features and the preset penalty coefficient as the exponent, an exponential penalty weight term is constructed; The fraud risk score is obtained by multiplying the exponential penalty weight term with the benchmark risk value, so that when semantic mutual exclusion between multimodal data leads to a sudden increase in cross-modal contradictory features, the fraud risk score increases exponentially.

10. A system for extracting and providing early warning of tensor features from multimodal data on pledged movable assets, employing the method for extracting and providing early warning of tensor features from multimodal data on pledged movable assets as described in any one of claims 1-9, characterized in that, include: Multimodal tensor construction module, tensor consensus decomposition module, modal residual extraction module, contradiction feature extraction module, and pledge risk early warning module; The multimodal tensor construction module is used to acquire multimodal data of movable property pledge business, which includes at least IoT sensing modal data and visual real-time modal data; and to perform feature mapping and time dimension alignment on the multimodal data to construct a multimodal high-order tensor containing time dimension and modal dimension. The tensor consensus decomposition module is used to perform low-rank tensor decomposition on multimodal high-order tensors, extract the consensus core tensor reflecting the cross-modal correlation characteristics between multi-source data and the factor matrix corresponding to each modality, and perform tensor multiplication operation on the consensus core tensor and the factor matrix along each modal dimension to reconstruct the consensus reconstruction tensor. The modal residual extraction module is used to calculate the difference between the multimodal high-order tensor and the consensus reconstruction tensor to obtain the global residual tensor, and to extract the IoT perception residual sub-tensor and the visual reality residual sub-tensor from the global residual tensor according to the modal dimension. The contradictory feature extraction module is used to calculate the cross-covariance between the IoT sensing residual sub-tensor and the visual reality residual sub-tensor, and construct a cross-modal interaction covariance tensor; perform feature decomposition on the cross-modal interaction covariance tensor, extract the feature value distribution spectrum, and calculate the dispersion index between the principal feature value and other feature values ​​in the feature value distribution spectrum, and determine the dispersion index as the cross-modal contradictory feature. The pledge risk early warning module is used to introduce cross-modal contradiction features and consensus core tensors into the risk scoring network and output fraud risk scores. When the fraud risk score exceeds a preset threshold, a movable property pledge risk warning is triggered; wherein, the risk scoring network maps the consensus core tensor to a benchmark risk value, and uses cross-modal contradiction features as a penalty adjustment factor to perform nonlinear penalty superposition on the benchmark risk value to obtain the fraud risk score.