A trade risk detection method, apparatus, device, medium and product
Patent Information
- Application Number
- CN202610448457.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-07
- Publication Date
- 2026-08-21
AI Technical Summary
[0004]本发明提供了一种贸易风险检测方法、装置、设备、介质及产品,以解决人工审核容易产生主观误判,以及无法识别跨时间周期的隐蔽违规交易的问题
[0010]本发明实施例的技术方案,响应于获取到电子贸易单证,将电子贸易单证转换为图像流,并将图像流输入至多模态大模型进行视觉语义对齐处理,得到贸易单证中的贸易信息,将贸易信息与风险数据库中的风险实体进行匹配,生成第一风险标签,基于贸易信息中的贸易主体,在审核日志数据库中查询贸易主体在设定时间窗口内的历史交易记录,并对历史交易记录进行时序风险检测,生成第二风险标签,基于所述第一风险标签和第二风险标签,确定针对所述电子贸易单证的贸易风险监测结果,通过在实体和时间序列维度分别进行风险检测,可以多维度识别贸易风险,精准检测到跨时间周期的隐蔽违规交易,并且通过多模态大模型的视觉语义对齐处理,可以精准提取贸易信息,避免主观误判。
Smart Images

Figure CN122618643A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of financial technology, and in particular to a method, apparatus, equipment, medium and product for detecting trade risks. Background Technology
[0002] With the digital transformation of global trade, commercial banks are experiencing exponential growth in the volume of cross-border remittances, letter of credit settlements, and trade finance transactions. To ensure trade security, compliance review of each transaction is crucial for risk control.
[0003] Currently, most commercial banks still heavily rely on manual review for compliance checks in their front-office operations when handling cross-border remittances, letter of credit verification, and other foreign exchange documentation. Manual review is prone to subjective errors and cannot identify covert irregularities that span multiple time periods. Summary of the Invention
[0004] This invention provides a trade risk detection method, apparatus, equipment, medium, and product to solve the problems of subjective misjudgment easily generated by manual review and the inability to identify hidden illegal transactions across time periods.
[0005] According to one aspect of the present invention, a trade risk detection method is provided, comprising: In response to obtaining electronic trade documents, the electronic trade documents are converted into an image stream, and the image stream is input into a multimodal large model for visual semantic alignment processing to obtain the trade information in the trade documents; The trade information is matched with risk entities in the risk database to generate a first risk label; Based on the trading entities in the trade information, the historical transaction records of the trading entities within a set time window are queried in the audit log database, and the historical transaction records are subjected to time-series risk detection to generate a second risk label; Based on the first risk label and the second risk label, the trade risk monitoring results for the electronic trade documents are determined.
[0006] According to another aspect of the present invention, a trade risk detection device is provided, comprising: The trade information recognition module is used to convert the electronic trade document into an image stream in response to the acquisition of the electronic trade document, and input the image stream into a multimodal large model for visual semantic alignment processing to obtain the trade information in the trade document. The first risk label generation module is used to match the trade information with risk entities in the risk database to generate a first risk label; The second risk label generation module is used to query the historical transaction records of the trading entity in the audit log database within a set time window based on the trading entity in the trade information, and to perform time-series risk detection on the historical transaction records to generate a second risk label. The risk detection result determination module is used to determine the trade risk monitoring result for the electronic trade document based on the first risk label and the second risk label.
[0007] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the trade risk detection method according to any embodiment of the present invention.
[0008] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the trade risk detection method according to any embodiment of the present invention.
[0009] According to another aspect of the present invention, a computer program product is provided, including a computer program that, when executed by a processor, implements the trade risk detection method of any embodiment of the present disclosure.
[0010] The technical solution of this invention, in response to obtaining electronic trade documents, converts the electronic trade documents into an image stream, and inputs the image stream into a multimodal large model for visual semantic alignment processing to obtain trade information in the trade documents. The trade information is then matched with risk entities in a risk database to generate a first risk label. Based on the trade entity in the trade information, the historical transaction records of the trade entity within a set time window are queried in the audit log database, and time-series risk detection is performed on the historical transaction records to generate a second risk label. Based on the first and second risk labels, the trade risk monitoring result for the electronic trade documents is determined. By performing risk detection in both entity and time-series dimensions, multi-dimensional trade risks can be identified, accurately detecting hidden illegal transactions across time periods. Furthermore, through the visual semantic alignment processing of the multimodal large model, trade information can be accurately extracted, avoiding subjective misjudgments.
[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a flowchart of a trade risk detection method provided in Embodiment 1 of the present invention; Figure 2 This is a flowchart of a trade risk detection method provided in Embodiment 2 of the present invention; Figure 3 This is a schematic diagram of the structure of a trade risk detection device according to Embodiment 3 of the present invention; Figure 4 This is a schematic diagram of the structure of an electronic device that implements the trade risk detection method of this invention. Detailed Implementation
[0014] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0015] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0016] Example 1 Figure 1The flowchart illustrates a trade risk detection method provided in Embodiment 1 of the present invention. This embodiment is applicable to situations where trade risks are detected at both the entity and time series dimensions. The method can be executed by a trade risk detection device, which can be implemented in hardware and / or software and can be configured in various general-purpose computing devices. Figure 1 As shown, the method includes: S110. In response to obtaining the electronic trade document, the electronic trade document is converted into an image stream, and the image stream is input into a multimodal large model for visual semantic alignment processing to obtain the trade information in the trade document.
[0017] Electronic trade documents are various unstructured electronic files involved in trade transactions, often with complex layouts, such as forms, seals, handwritten signatures, and blurry scanned information. Examples of electronic trade documents include Portable Document Format (PDF), Joint Photographic Experts Group (JPG), and Portable Network Graphics (PNG).
[0018] Image streams are created by converting the original electronic trade documents into a high-resolution, continuous sequence of pixel data using a rendering engine. The image streams retain the complete layout information of the original electronic trade documents, including table borders, seal locations, and visual features such as handwriting.
[0019] Multimodal large models are deep learning models capable of processing both image and text data simultaneously. They are acquired through pre-training on large amounts of image and text data to learn cross-modal semantic alignment capabilities (i.e., visual semantic alignment capabilities).
[0020] Trade information is a structured field extracted from electronic trade documents. For example, trade information includes the trading entity (including the buyer's name and the seller's name), transaction amount, currency, goods description, and date.
[0021] In this embodiment of the invention, after acquiring the electronic trade documents, a rendering engine converts them into an image stream. This preserves the layout information of the electronic trade documents while standardizing the data format of documents with different formats, providing a standardized input for subsequent multimodal processing of electronic documents. The converted image stream is stored in pixel matrix form, retaining visual features such as table borders, seal positions, and handwriting in the electronic trade documents. Regardless of whether the original document is a PDF, scanned copy, or photograph, it can be converted into a unified image stream, providing standard input for subsequent processing and preserving specific layout information as a basis for accurate extraction of trade information.
[0022] Furthermore, the converted image stream is input into a multimodal large model for visual semantic alignment processing, establishing the correlation between semantic features and image features in electronic trade documents, and ultimately identifying trade information in electronic trade documents. Visual semantic alignment processing can improve the accuracy of trade information recognition.
[0023] Optional electronic trade documents may be commercial invoices, bills of lading, or customs declarations; the format of electronic trade documents may be portable document format, portable web image, or joint image expert group.
[0024] In this optional embodiment, the types and specific formats of electronic trade documents are provided, wherein electronic trade documents may be commercial invoices, bills of lading or customs declarations, etc.; electronic trade documents may be in portable document formats, portable web images or Joint Image Experts Group formats, etc.
[0025] S120. Match trade information with risk entities in the risk database to generate the first risk label.
[0026] The risk database is used to store various risk entities and their attributes that require risk monitoring. Among them, a risk entity is a specific object to be monitored, such as the trade address or goods description of a sensitive trading entity with potential trade risks. Each risk entity can be associated with a risk type (e.g., sanctions and split transactions) and a risk level (e.g., low risk, medium risk, and high risk).
[0027] The first risk label is a risk indication generated based on the matching results of current trade information and risk database. The first risk label may include the risk entity, the type of risk associated with the risk entity, and the risk level.
[0028] In this embodiment of the invention, a field to be detected is extracted sequentially from trade information. For example, the field to be detected may be the trading entity (buyer's name, seller's name), delivery address, or goods description. The extracted fields are preprocessed, for example, by standardizing capitalization and removing special characters, to improve matching accuracy. Furthermore, the preprocessed fields are compared with risk entities in a risk database. If a field successfully matches a risk entity in the risk database, a first risk label is generated. The first risk label may include the identifier of the successfully matched risk entity, the corresponding risk type, and the risk level. By identifying trade entities and matching them with risk entities, the detection efficiency and accuracy are significantly improved compared to manual comparison.
[0029] In a specific example, the goods description information is extracted from the trade information as the field to be detected. This field is then matched with the risk entities (sanctioned entities) of the goods description type in the risk database. If a match is successful with the target risk entity, a first risk label is generated, which includes the risk entity's identifier, the risk type (goods embargo risk) corresponding to the risk entity, and the risk level (high risk).
[0030] S130. Based on the trade entities in the trade information, query the historical transaction records of the trade entities in the audit log database within a set time window, and perform time-series risk detection on the historical transaction records to generate a second risk label.
[0031] The trading entities refer to the two parties involved in the trade, including the buyer's name and the seller's name; the audit log database is a database that persistently stores historical audit records, in which each record may include information such as audit time, trading entity, transaction amount, risk label, and reasoning logic.
[0032] The second risk label is a risk indicator generated based on the time-series risk detection results, used to mark whether the current transaction has risks at the level of historical behavior patterns.
[0033] In this embodiment of the invention, to detect hidden irregularities that cannot be detected by a single electronic trade document, such as the behavior of circumventing large-scale fund supervision through multiple small transactions, a time dimension is introduced for in-depth defense. Specifically, based on the trade entity in the extracted trade information, such as the buyer's name, the historical transaction records of the trade entity within a preset time window, such as the past 3 days, are queried in the audit log database, and time-series risk detection is performed to generate a second risk label. Specifically, the retrieved historical transaction records are statistically analyzed to determine whether a preset abnormal transaction pattern exists. Taking split transactions as an example, the detection logic is as follows: count the number of transactions with a single amount less than a first threshold within the time window; if the number reaches or exceeds a second threshold, it is determined that there is a high-frequency small-amount transaction characteristic. Further, it is determined whether the amount of the current transaction is also less than the first threshold; if so, it is confirmed that the current transaction also belongs to the abnormal pattern.
[0034] If an abnormal pattern is detected, a corresponding secondary risk label is generated, such as "Split Transaction Risk - High". Furthermore, other time-series patterns can be added, such as sudden increases in amount within a short period or abnormal overlap with historical counterparties. If no abnormalities are found, a "No Historical Behavior Risk" label is generated. By using time-series risk detection, individual trades can be identified as compliant, but combined with historical records, trade risks related to illegal trading activities can be discovered.
[0035] S140. Based on the first risk label and the second risk label, determine the trade risk monitoring results for electronic trade documents.
[0036] In this embodiment of the invention, the generated first risk label and second risk label are fused, and the final trade risk monitoring result is determined according to a preset decision logic. Specifically, the first risk label and the second risk label are obtained, and each label may contain a risk type and risk level (such as high, medium, and low). A comprehensive judgment is made according to preset fusion rules. For example, if any label indicates high risk (such as "sanction risk - high" or "split transaction risk - high"), a high-risk trade risk detection result is directly output; if there is a medium-risk label, a trade risk detection result of "requires manual review" is output; if all labels are risk-free or low-risk, a "pass" trade risk detection result is output.
[0037] The technical solution of this invention, in response to obtaining electronic trade documents, converts the electronic trade documents into an image stream, and inputs the image stream into a multimodal large model for visual semantic alignment processing to obtain trade information in the trade documents. The trade information is then matched with risk entities in a risk database to generate a first risk label. Based on the trade entity in the trade information, the historical transaction records of the trade entity within a set time window are queried in the audit log database, and time-series risk detection is performed on the historical transaction records to generate a second risk label. Based on the first and second risk labels, the trade risk monitoring results for the electronic trade documents are determined. By performing risk detection in both entity and time-series dimensions, multi-dimensional trade risks can be identified, accurately detecting hidden illegal transactions across time periods. Furthermore, through the visual semantic alignment processing of the multimodal large model, trade information can be accurately extracted, avoiding subjective misjudgments.
[0038] Example 2 Figure 2 This is a flowchart of a trade risk detection method provided in Embodiment 2 of the present invention. This embodiment further refines the above embodiment, providing specific steps for inputting an image stream into a multimodal large model for visual semantic alignment processing to obtain trade information in trade documents, and specific steps for performing time-series risk detection on historical transaction records to generate a second risk label. Figure 2 As shown, the method includes: S210. In response to obtaining electronic trade documents, the electronic trade documents are converted into an image stream. The image stream is input into a multimodal large model. The multimodal large model performs block encoding on the image stream to determine the visual features of the electronic trade documents and performs semantic encoding on the text fragments in the image stream to generate text features.
[0039] In this embodiment of the invention, the system first receives electronic trade documents uploaded by the user. These documents may be in a portable document format, a portable web image format, or a Joint Picture Experts Group (JPI) format. To eliminate the heterogeneity caused by different formats and preserve layout details, a rendering engine is invoked to convert the original file into a high-resolution image stream. The converted image stream is stored in the form of a pixel matrix, fully preserving visual features such as table borders, seal positions, and handwritten strokes.
[0040] Furthermore, the image stream is input into a multimodal large model, which uses a visual encoder to segment the image stream into fixed-size image blocks. Each image block is then subjected to linear projection and positional encoding, followed by feature extraction, ultimately generating the visual features of the electronic trade document.
[0041] In a specific example, the electronic trade document uploaded by the user is a commercial invoice. Through block encoding, image blocks can be obtained for the stamp area, the handwritten signature area, the printed text area, and the table border area. The visual features of the stamp area image block include texture information such as red color, rounded edges, and semi-transparency; the visual features of the handwritten signature area image block include handwriting characteristics such as continuous strokes and uneven thickness; and the visual features of the printed text area image block include characteristics such as regular font shape and neat arrangement.
[0042] Meanwhile, the multimodal model, through its built-in text recognition and encoding module, detects text regions from the image stream and identifies continuous text belonging to the same semantic unit as text segments. Each text segment is then processed by a text encoder to generate text features.
[0043] S220. Through the cross-attention mechanism, establish the correspondence between visual features and text features, and based on the correspondence, aggregate the visual features into the text features to obtain enhanced text features that fuse visual features.
[0044] In this embodiment of the invention, a correspondence between visual and text features is established through a cross-attention mechanism, and the visual features are aggregated into the text features to generate enhanced text features. Specifically, using text features as queries and visual features as keys and values, the attention weights of each text fragment with all image blocks are calculated through the cross-attention mechanism, and the visual features are weighted and aggregated onto the text features to obtain enhanced text features that fuse visual information. This ensures that each text fragment not only contains its own semantics but also carries the visual attributes of its region, such as seal texture or handwritten strokes.
[0045] In a specific example, a text fragment located below a seal can serve as a query. Through cross-attention, the most attention-grabbing image blocks are determined to be the text area below the seal (although partially obscured, some strokes are still visible) and the surrounding red seal area. After aggregation, the enhanced features of this text fragment include visual attributes such as "red," "circular edges," and "semi-transparent occlusion."
[0046] S230. Contextual modeling of enhanced text features is performed through a self-attention mechanism, and multiple regions to be identified in the image stream are determined based on the correspondence between visual features and text features; the regions to be identified include the stamp occlusion region, the handwritten signature region, and the printed text region.
[0047] In this embodiment of the invention, based on the correspondence between visual features and text features established through enhanced text features and the cross-attention stage, global context modeling is performed through a self-attention mechanism to determine multiple regions to be identified in the image stream. Specifically, enhanced text features are input into the self-attention module. In the self-attention calculation, each text segment serves as both a query and a key / value pair, calculating its attention weight with other text segments to achieve information interaction and obtain feature vectors of text segments that incorporate global context information.
[0048] Furthermore, during the self-attention modeling process, the model simultaneously references the correspondence between each text fragment and an image patch. Through this association, the model can infer the spatial proximity between text fragments. If two text fragments both highly focus on the same image region, then these two text fragments are spatially adjacent and belong to the same region.
[0049] Based on feature vectors that integrate global context and spatial proximity, the model clusters text fragments, grouping spatially adjacent fragments with similar features (such as both possessing the visual attributes of a seal) into the same region to be identified. Simultaneously, based on the visual attributes carried by each text fragment—for example, a red circular texture corresponding to a seal-occluded area, illegible strokes corresponding to a handwritten signature area, and regular characters corresponding to printed text areas—a type label is assigned to each region.
[0050] For each region to be identified obtained from clustering, the model determines the boundary of the region in the image stream based on the image block position corresponding to the text fragments it contains, providing spatial guidance for subsequent key field localization.
[0051] S240. Locate the target area to be identified for each field in the trade information, extract the text content from the target area to be identified, and output the trade information.
[0052] In this embodiment of the invention, based on the pre-determined region to be identified, the target region to be identified for each field in the trade information is located, and text content is extracted from the target region to finally output structured trade information. Specifically, according to preset association rules between field types and region types, the possible region type for each field is determined. For example, core fields such as buyer's name, seller's name, amount, and date typically appear in the printed text area.
[0053] Furthermore, for each field, such as "buyer's name," regions matching the type are filtered from the identified areas to be identified, such as all printed text areas. Based on layout rules, such as the buyer's name typically being in the upper left corner of the invoice, the candidate range is further narrowed down to determine the candidate regions. Further, within the candidate regions, the specific location of the field is located using a detection head. For example, the detection head can combine the features of text fragments within the region with their spatial location to output the bounding box coordinates of the field.
[0054] Furthermore, for the located field areas, an appropriate decoding strategy is used to extract the text content based on the area type. For example, for printed text areas, conventional decoding is used to read the clear printed text; for areas obscured by seals, visual cues carried in enhanced text features are used, combined with an attention mechanism, to complete and infer the obscured parts, extracting the complete text content; for handwritten signature areas, a handwriting recognition decoder is used to extract the handwritten text or signature image. Finally, the extracted field content is organized according to a predefined format, outputting standardized trade information. By establishing the correlation between text features and visual features, and enabling the text to be understood in a global semantic context, the accuracy of area determination is improved, thereby enhancing the accuracy of content recognition in electronic trade documents.
[0055] S250: Match trade information with risk entities in the risk database to generate the first risk label.
[0056] Optionally, the trade information is matched with risk entities in the risk database to generate a first risk label, including: One field is extracted sequentially from the trade information as the field to be matched, and the field to be matched is embedded to obtain the vector to be matched; The similarity of the vector to be matched with the entity vector corresponding to the risk entity in the risk database is compared. If the similarity between a target entity vector and the vector to be matched is less than a set similarity threshold in the risk database, a first risk label containing the risk entity identifier associated with the target entity vector is generated.
[0057] In this optional embodiment, a specific method is provided for matching trade information with risk entities in a risk database to generate a first risk label: A field is sequentially extracted from the trade information as a field to be matched, and this field is embedded to obtain a corresponding vector to be matched. Further, the vector to be matched is compared with the entity vectors corresponding to the risk entities in the risk database for similarity. If a target entity vector in the risk database has a similarity less than a set similarity threshold, a first risk label containing the risk entity identifier associated with the target entity vector is generated. The first risk label may also include the risk type and risk level corresponding to the risk entity identifier. By using vector comparison, the need for manual verification of each field in the trade information is eliminated, improving the efficiency and accuracy of trade risk screening.
[0058] S260. Based on the trade entities in the trade information, query the historical transaction records of the trade entities in the audit log database within a set time window, and query the number of historical trades in which the single trade amount is less than the amount threshold.
[0059] In this embodiment of the invention, based on the trading entity in the trade information, such as a buyer or seller, the historical transaction records of the trading entity within a set time window are queried in the audit log database, for example, historical transaction records within the past 3 days. Furthermore, in the historical transaction records, the number of historical trades where the single trade amount is less than a threshold is queried to determine whether small-amount, high-frequency transactions exist.
[0060] S270. If the number of historical trades exceeds a set threshold, the trade amount in the electronic trade document is compared with the threshold amount.
[0061] In this embodiment of the invention, if the number of historical trades exceeds a set threshold, it indicates that there are small-amount high-frequency transactions. At this time, the trade amount in the electronic trade document is compared with the amount threshold to determine whether the current trade belongs to the aforementioned small-amount high-frequency transactions.
[0062] S280. If the trade amount is less than the amount threshold, generate a second risk label containing the risk identifier of split transaction.
[0063] In this embodiment of the invention, when the trade amount is less than a threshold, a second risk label containing a risk identifier for splitting the transaction is generated. This second risk label also includes a risk level. By transforming discrete document review into continuous behavioral pattern analysis, hidden compliance risks across documents and time periods are effectively identified. This avoids the problem that manual review typically relies on static rule judgments based on single transactions, lacking the ability to perform dynamic correlation analysis across time dimensions.
[0064] S290. Based on the first risk label and the second risk label, determine the trade risk monitoring results for electronic trade documents.
[0065] Optionally, based on the first risk label and the second risk label, the trade risk monitoring results for electronic trade documents are determined, including: If at least one of the first risk label and the second risk label contains a risk indicator, the trade risk detection result of the electronic trade document is determined to be risky trade.
[0066] In this optional embodiment, a specific method is provided for determining the trade risk monitoring result for electronic trade documents based on a first risk label and a second risk label: the contents of the first and second risk labels are read sequentially. If the first risk label contains a risk entity identifier, and / or the second risk label contains a split transaction risk identifier, then the trade risk detection result of the current electronic trade document is determined to be a risky trade. Subsequently, for risky trade, a risk warning can be initiated, and the risk warning information can include the risk identifier content from the first and second risk labels. The risk entities in the risk database are dynamically updated according to actual risk control and compliance requirements.
[0067] By comparing the first risk label generated with a risk database and the second risk label generated by time-series risk detection, the final risk detection result can be determined, which can identify trade risks in different dimensions and improve the accuracy and reliability of risk detection.
[0068] Optionally, after determining the trade risk monitoring results for electronic trade documents based on the first and second risk labels, the following may also be included: Calculate the hash value of electronic trade documents; Trade information, first risk label, second risk label, trade risk detection results, hash value, and inference data of multimodal large model are persistently stored in the audit log database according to a predefined format; In response to a user's export request, the system serializes the data stored in the audit log database and generates a downloadable standard format file.
[0069] In this optional embodiment, to achieve full-process data traceability, the hash value of the electronic trade document is calculated. Then, the trade information, first risk label, second risk label, trade risk detection result, hash value, and inference data from the multimodal large model are persistently stored in the audit log table of the audit log database according to a predefined format. Specifically, each data entry in the audit log table includes a primary key identifier, audit time, trade information (buyer's name, seller's name, and amount), first risk label, second risk label, trade risk detection result, hash value, and inference data from the multimodal large model. This achieves full lifecycle management of data; each audit is transformed into a structured digital asset, which not only meets the regulatory agency's rigid requirements for audit traceability but also provides labeled data for subsequent optimization of the risk control model.
[0070] Furthermore, to activate the static data in the audit log database, a serialization engine is deployed in the export interface. When a user terminal initiates an export request, the export interface performs a full table scan or conditional search on the audit log table, performs character set adaptation and serialization conversion, and exports millions of unstructured log entries into standard delimited value files or tabular files with one click, allowing compliance departments to generate regulatory reports. The data serialization export function reduces the time-consuming manual ledger statistics to seconds, providing high-quality labeled data feedback for iterative risk control models.
[0071] The technical solution of this invention inputs the image stream of electronic trade documents into a multimodal large model for visual semantic alignment processing to obtain trade information in the trade documents. By establishing the correlation between text features and visual features, and enabling the text to be understood in a global semantic context, the accuracy of region determination is improved, thereby improving the accuracy of content recognition in electronic trade documents, thus improving the reliability of risk detection, avoiding false alarms and missed alarms. Furthermore, by performing time-series risk detection on the historical transaction records of trade entities, hidden illegal transactions across time periods can be identified.
[0072] Example 3 Figure 3 This is a schematic diagram of a trade risk detection device provided in Embodiment 3 of the present invention. Figure 3 As shown, the device includes: The trade information recognition module 310 is used to convert the electronic trade document into an image stream in response to obtaining the electronic trade document, and input the image stream into a multimodal large model for visual semantic alignment processing to obtain the trade information in the trade document. The first risk label generation module 320 is used to match the trade information with risk entities in the risk database to generate a first risk label; The second risk label generation module 330 is used to query the historical transaction records of the trading entity in the audit log database within a set time window based on the trading entity in the trade information, and to perform time-series risk detection on the historical transaction records to generate a second risk label. The risk detection result determination module 340 is used to determine the trade risk monitoring result for the electronic trade document based on the first risk label and the second risk label.
[0073] The technical solution of this invention, in response to obtaining electronic trade documents, converts the electronic trade documents into an image stream, and inputs the image stream into a multimodal large model for visual semantic alignment processing to obtain trade information in the trade documents. The trade information is then matched with risk entities in a risk database to generate a first risk label. Based on the trade entity in the trade information, the historical transaction records of the trade entity within a set time window are queried in the audit log database, and time-series risk detection is performed on the historical transaction records to generate a second risk label. Based on the first and second risk labels, the trade risk monitoring result for the electronic trade documents is determined. By performing risk detection in both entity and time-series dimensions, multi-dimensional trade risks can be identified, accurately detecting hidden illegal transactions across time periods. Furthermore, through the visual semantic alignment processing of the multimodal large model, trade information can be accurately extracted, avoiding subjective misjudgments.
[0074] Optional, the trade information identification module 310 is specifically used for: The image stream is input into a multimodal large model, which performs block encoding on the image stream to determine the visual features of the electronic trade document, and performs semantic encoding on the text fragments in the image stream to generate text features; By using a cross-attention mechanism, a correspondence between the visual features and the text features is established, and based on the correspondence, the visual features are aggregated into the text features to obtain enhanced text features that fuse visual features. The enhanced text features are modeled in context using a self-attention mechanism, and multiple regions to be identified in the image stream are determined based on the correspondence between the visual features and the text features. The regions to be identified include areas covered by a stamp, areas of handwritten signatures, and areas of printed text. Locate the target region to be identified for each field in the trade information, extract the text content from the target region to be identified, and output the trade information.
[0075] Optionally, the second risk label generation module 330 is specifically used for: In the historical transaction records, query the number of historical transactions where the amount of a single transaction is less than the amount threshold; If the number of historical trades exceeds a set threshold, the trade amount in the electronic trade document will be compared with the threshold amount. If the trade amount is less than the amount threshold, a second risk label containing a split transaction risk identifier is generated.
[0076] Optional trade risk detection devices also include: The hash value calculation module is used to calculate the hash value of the electronic trade document after determining the trade risk monitoring result for the electronic trade document based on the first risk label and the second risk label. The data storage module is used to persistently store the trade information, the first risk label, the second risk label, the trade risk detection result, the hash value, and the inference data of the multimodal large model into the audit log database in a predefined format. The file export module is used to respond to the user's export request, perform serialization conversion on the data stored in the audit log database, and generate a downloadable standard format file.
[0077] Optionally, the first risk label generation module 320 is specifically used for: One field is extracted sequentially from the trade information as the field to be matched, and the field to be matched is embedded to obtain the vector to be matched; The similarity of the vector to be matched with the entity vector corresponding to the risk entity in the risk database is compared. If the similarity between a target entity vector and the vector to be matched is less than a set similarity threshold in the risk database, a first risk label containing the risk entity identifier associated with the target entity vector is generated.
[0078] Optionally, the risk detection result determination module 340 is specifically used for: If at least one of the first risk label and the second risk label contains a risk identifier, the trade risk detection result of the electronic trade document is determined to be a risky trade.
[0079] Optionally, the electronic trade document is a commercial invoice, bill of lading, or customs declaration; the format of the electronic trade document is a portable document format, portable web image, or joint image expert group.
[0080] The trade risk detection device provided in this embodiment of the invention can execute the trade risk detection method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.
[0081] In the technical solution of this invention, the information collected is information and data authorized by the user or fully authorized by all parties. The collection, storage, use, processing, transmission, provision, disclosure and application of related data all comply with the relevant laws, regulations and standards of relevant countries and regions, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entry points for users to choose to authorize or refuse.
[0082] Example 4 According to embodiments of the present invention, the present invention also provides an electronic device, a readable storage medium, and a computer program product.
[0083] Figure 4 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, application processors, blade application processors, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0084] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory 12 or a random access memory 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the read-only memory 12 or a computer program loaded from storage unit 18 into the random access memory 13. The random access memory 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, read-only memory 12, and random access memory 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0085] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0086] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, central processing units, graphics processing units, various special-purpose artificial intelligence computing chips, various processors running machine learning model algorithms, digital signal processors, and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as trade risk detection methods.
[0087] In some embodiments, the trade risk detection method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via read-only memory 12 and / or communication unit 19. When the computer program is loaded into random access memory 13 and executed by processor 11, one or more steps of the trade risk detection method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the trade risk detection method by any other suitable means (e.g., by means of firmware).
[0088] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays, application-specific integrated circuits (ASICs), application-specific standard products (ASICs), systems-on-a-chip (SoCs), complex programmable logic devices, computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0089] Computer programs used to implement the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs can be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or application.
[0090] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, optical fibers, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0091] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a cathode ray tube or liquid crystal display monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0092] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data application processors), or computing systems that include middleware components (e.g., application application processors), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0093] A computing system can include clients and applications. Clients and applications are generally geographically separated and typically interact via communication networks. The client-application relationship is established by computer programs running on the respective computers and having a client-application relationship with each other. An application can be a cloud application, also known as a cloud computing application or cloud host, which is a host product within the cloud computing application architecture to address the shortcomings of traditional physical hosts and virtual private services, such as high management difficulty and weak business scalability.
[0094] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0095] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for detecting trade risks, characterized in that, include: In response to obtaining electronic trade documents, the electronic trade documents are converted into an image stream, and the image stream is input into a multimodal large model for visual semantic alignment processing to obtain the trade information in the trade documents; The trade information is matched with risk entities in the risk database to generate a first risk label; Based on the trading entities in the trade information, the historical transaction records of the trading entities within a set time window are queried in the audit log database, and the historical transaction records are subjected to time-series risk detection to generate a second risk label; Based on the first risk label and the second risk label, the trade risk monitoring results for the electronic trade documents are determined.
2. The method according to claim 1, characterized in that, The image stream is input into a multimodal large model for visual semantic alignment processing to obtain the trade information in the trade documents, including: The image stream is input into a multimodal large model, which performs block encoding on the image stream to determine the visual features of the electronic trade document, and performs semantic encoding on the text fragments in the image stream to generate text features; By using a cross-attention mechanism, a correspondence between the visual features and the text features is established, and based on the correspondence, the visual features are aggregated into the text features to obtain enhanced text features that fuse visual features. The enhanced text features are modeled in context using a self-attention mechanism, and multiple regions to be identified in the image stream are determined based on the correspondence between the visual features and the text features. The regions to be identified include areas covered by a stamp, areas of handwritten signatures, and areas of printed text. Locate the target region to be identified for each field in the trade information, extract the text content from the target region to be identified, and output the trade information.
3. The method according to claim 1, characterized in that, Perform time-series risk detection on the historical transaction records to generate a second risk label, including: In the historical transaction records, query the number of historical transactions where the amount of a single transaction is less than the amount threshold; If the number of historical trades exceeds a set threshold, the trade amount in the electronic trade document will be compared with the threshold amount. If the trade amount is less than the amount threshold, a second risk label containing a split transaction risk identifier is generated.
4. The method according to claim 1, characterized in that, After determining the trade risk monitoring results for the electronic trade document based on the first risk label and the second risk label, the method further includes: Calculate the hash value of the electronic trade document; The trade information, the first risk label, the second risk label, the trade risk detection result, the hash value, and the inference data of the multimodal large model are persistently stored in the audit log database according to a predefined format. In response to a user's export request, the data stored in the audit log database is serialized and converted to generate a downloadable standard format file.
5. The method according to claim 1, characterized in that, The trade information is matched with risk entities in the risk database to generate a first risk label, including: One field is extracted sequentially from the trade information as the field to be matched, and the field to be matched is embedded to obtain the vector to be matched; The similarity of the vector to be matched with the entity vector corresponding to the risk entity in the risk database is compared. If the similarity between a target entity vector and the vector to be matched is less than a set similarity threshold in the risk database, a first risk label containing the risk entity identifier associated with the target entity vector is generated.
6. The method according to claim 1, characterized in that, Based on the first risk label and the second risk label, the trade risk monitoring results for the electronic trade documents are determined, including: If at least one of the first risk label and the second risk label contains a risk identifier, the trade risk detection result of the electronic trade document is determined to be a risky trade.
7. The method according to any one of claims 1-6, characterized in that, The electronic trade documents are commercial invoices, bills of lading, or customs declarations; the format of the electronic trade documents is portable document format, portable web image, or joint image expert group.
8. A trade risk detection device, characterized in that, include: The trade information recognition module is used to convert the electronic trade document into an image stream in response to the acquisition of the electronic trade document, and input the image stream into a multimodal large model for visual semantic alignment processing to obtain the trade information in the trade document. The first risk label generation module is used to match the trade information with risk entities in the risk database to generate a first risk label; The second risk label generation module is used to query the historical transaction records of the trading entity in the audit log database within a set time window based on the trading entity in the trade information, and to perform time-series risk detection on the historical transaction records to generate a second risk label. The risk detection result determination module is used to determine the trade risk monitoring result for the electronic trade document based on the first risk label and the second risk label.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the trade risk detection method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the trade risk detection method according to any one of claims 1-7.
11. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the trade risk detection method according to any one of claims 1-7.