A Business Risk Identification Method and System Based on Container Shipping Documents

By extracting feature, deeply mining and fusion of sea shipping documents, combined with multiple risk identification methods, the traditional manual auditing is solved, and more accurate and reliable business risk identification is achieved.

CN119515089BActive Publication Date: 2025-06-24SHANGHAI COSCO SHIPPING CONTAINER TRANSPORT INFORMATION SERVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510088663.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-06-24
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

The traditional manual audit method is inefficient and error-prone in the audit of container shipping documents in the maritime field, making it difficult to effectively identify potential business risks.

Method used

By obtaining the document information to be verified related to the target container, performing feature extraction processing, conducting in-depth feature mining at multiple levels, feature fusion, and multi-faceted risk identification based on the fusion feature vector.

Benefits of technology

It improves the accuracy and reliability of business risk identification, and efficient review of maritime documents is achieved through an automated system, reducing the possibility of manual errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119515089B_ABST
    Figure CN119515089B_ABST
Patent Text Reader

Abstract

An embodiment of this specification discloses a method and system for identifying business risks based on container shipping documents, which relates to the technical fields of artificial intelligence and logistics transportation. Among them, the method includes: obtaining the information of the documents to be verified related to the target container; performing feature extraction processing on the status information and / or content information of each sub-item in the information of the documents to be verified to obtain an initial feature vector; performing deep feature mining at multiple different levels on the initial feature vector to obtain multiple deep feature vectors corresponding to the information of the documents to be verified, where each deep feature vector is used to represent the deep features of the information of the documents to be verified at the corresponding level; performing feature fusion on the multiple deep feature vectors corresponding to the information of the documents to be verified to obtain a fused feature vector; performing risk identification in multiple aspects based on the fused feature vector to obtain a corresponding business risk identification result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical fields of artificial intelligence and logistics transportation. Specifically, it relates to a method and system for identifying business risks based on container shipping documents. Background Art

[0002] In the field of maritime transportation, container shipping documents (such as booking notes, bills of lading, packing lists, cargo manifests, endorsement lists, customs declarations, commercial invoices, etc.) carry various information during the process of cargo transportation. These documents are not only important vouchers for cargo transportation but also important bases for identifying business risks.

[0003] There may be various risk information contained in shipping documents, such as false information, illegal trade, compliance issues, unclear cargo transportation routes, etc. With the increase in global trade and the improvement of complexity, the review and risk identification of documents have become increasingly important. However, the traditional manual review method is inefficient and prone to errors (such as missed identification). Therefore, there is an urgent need for an automated system to implement the review of container shipping documents to identify potential business risks. Summary of the Invention

[0004] In view of this, one aspect of the embodiments of this specification provides a method for identifying business risks based on container shipping documents, and the method includes:

[0005] Obtain the information of the documents to be verified related to the target container, where the information of the documents to be verified is obtained by identifying at least one document related to the target container and arranging them in a preset order;

[0006] Extract feature processing on the status information and / or content information of each sub-item in the information of the documents to be verified to obtain an initial feature vector;

[0007] Perform deep feature mining at multiple different levels on the initial feature vector to obtain multiple deep feature vectors corresponding to the information of the documents to be verified, where each deep feature vector is used to represent the deep feature of the information of the documents to be verified at the corresponding level;

[0008] Fuse the features of the multiple deep feature vectors corresponding to the information of the documents to be verified to obtain a fused feature vector;

[0009] Based on the fused feature vector, perform risk identification in multiple aspects to obtain the corresponding business risk identification result.

[0010] In some embodiments, the obtaining of the information of the documents to be verified related to the target container includes:

[0011] For each target document related to the target container, the document type information of the target document and the status information and / or content information of at least one sub-item included in the target document are extracted through OCR recognition technology, and the partial information to be verified corresponding to the target document is obtained;

[0012] The status information and / or content information of at least one sub-item included in each of the target documents are arranged in a first preset order, and when the number of the documents is multiple, the partial information to be verified corresponding to the multiple documents is arranged in a second preset order to obtain the information of the documents to be verified related to the target container.

[0013] In some embodiments, the step of arranging the partial information to be verified corresponding to the multiple documents in a second preset order to obtain the information of the documents to be verified related to the target container includes:

[0014] Sort according to the importance of each document to determine the second preset order;

[0015] Write the corresponding partial information to be verified into the corresponding position in the second preset order according to the document type information of at least one document related to the target container, and fill in the missing data with a preset value to obtain the information of the documents to be verified related to the target container.

[0016] In some embodiments, the step of performing feature extraction processing on the status information and / or content information of each sub-item in the information of the documents to be verified to obtain an initial feature vector includes:

[0017] Perform semantic encoding on the status information and / or content information of each sub-item in the information of the documents to be verified through natural language processing to obtain at least one semantic vector corresponding to the information of the documents to be verified, and obtain the initial feature vector based on the at least one semantic vector.

[0018] In some embodiments, the step of performing deep feature mining on the initial feature vector at multiple different levels to obtain multiple deep feature vectors corresponding to the information of the documents to be verified includes:

[0019] Input the initial feature vector into a deep feature extraction network, where the deep feature extraction network includes multiple hidden layers, each hidden layer corresponds to an activation function, and each of the hidden layers and the activation function is used to perform one-level deep feature mining on the initial feature vector;

[0020] Based on the outputs of the multiple hidden layers, obtain multiple deep feature vectors corresponding to the information of the documents to be verified.

[0021] In some embodiments, the performing feature fusion on multiple deep feature vectors corresponding to the document information to be verified to obtain the fused feature vector includes: concatenating the multiple deep feature vectors to obtain the fused feature vector.

[0022] In some embodiments, the risk identification in multiple aspects is performed based on the fused feature vector to obtain corresponding business risk identification results, including:

[0023] The fused feature vectors are respectively input into a plurality of pre-trained classifiers for classification processing, and the business risk identification result is obtained based on the outputs of the plurality of classifiers; wherein each of the classifiers corresponds to one aspect of risk identification.

[0024] In some embodiments, the multiple classifiers are trained based on the following method:

[0025] Acquire a sample fusion feature vector obtained by processing at least one sample document related to the sample case, and a plurality of risk labels corresponding to the sample fusion feature vector, wherein each of the plurality of risk labels is used as an expected output of a classifier;

[0026] Inputting the sample fusion feature vector into the multiple classifiers respectively to obtain the prediction output corresponding to each classifier;

[0027] Determine a loss value for each classifier based on a difference between a predicted output and an expected output corresponding to each of the classifiers;

[0028] The multiple classifiers are iteratively optimized based on the loss values ​​until the multiple trained classifiers are obtained when preset training conditions are met.

[0029] In some embodiments, before inputting the sample fusion feature vector into the multiple classifiers respectively to obtain the prediction output corresponding to each classifier, the method further includes:

[0030] For each of the classifiers, the sample fusion feature vector is attention-encoded by an attention mechanism to obtain an attention fusion vector corresponding to each classifier;

[0031] The attention fusion vector is input as the input data of the classifier into the corresponding classifier for processing to obtain the corresponding prediction output of each classifier.

[0032] Another aspect of the embodiments of this specification further provides a business risk identification system based on container shipping documents, the system comprising:

[0033] An acquisition module, configured to acquire the document information to be verified related to the target container, where the document information to be verified is obtained by identifying at least one document related to the target container and arranging them in a preset order;

[0034] An initial feature vector determination module, configured to perform feature extraction processing on the status information and / or content information of each sub-item in the document information to be verified, to obtain an initial feature vector;

[0035] A deep feature vector determination module, configured to perform deep feature mining at multiple different levels on the initial feature vector, to obtain multiple deep feature vectors corresponding to the document information to be verified, where each deep feature vector is used to represent the deep feature of the document information to be verified at the corresponding level;

[0036] A fused feature vector determination module, configured to perform feature fusion on the multiple deep feature vectors corresponding to the document information to be verified, to obtain a fused feature vector;

[0037] A business risk identification module, configured to perform risk identification in multiple aspects based on the fused feature vector, to obtain a corresponding business risk identification result.

[0038] The beneficial effects that the business risk identification method and system based on container shipping documents provided by the embodiments of the present specification may bring at least include: By performing feature extraction processing on the document information to be verified to obtain an initial feature vector, and then performing deep feature mining at multiple different levels on the initial feature vector to obtain multiple deep feature vectors, the complex features in the document information to be verified can be more comprehensively reflected, thereby improving the accuracy and reliability of the business risk identification result obtained based on the multiple deep feature vectors in the subsequent process.

[0039] Additional features will be partially described below. For those skilled in the art, it will become obvious by referring to the following content and drawings, or can be understood by generating or operating examples. The features of the present specification can be realized and obtained by practicing or using various aspects of the methods, tools, and combinations described in the following detailed examples. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] The present specification will be further described in the form of exemplary embodiments, and these exemplary embodiments will be described in detail through the drawings. These embodiments are not restrictive. In these embodiments, the same numbers represent the same structures, where:

[0041] Figure 1 is an exemplary flowchart of a business risk identification method based on container shipping documents shown in some embodiments of the present specification;

[0042] Figure 2 It is a schematic diagram of a feature vector processing flow shown in some embodiments of this specification;

[0043] Figure 3 It is a schematic diagram of the principle of business risk identification shown in some embodiments of this specification;

[0044] Figure 4 It is a schematic diagram of the principle of business risk identification shown in some other embodiments of this specification;

[0045] Figure 5 It is an exemplary module diagram of a business risk identification system based on container shipping documents shown in some embodiments of this specification. Detailed implementation manners

[0046] To more clearly illustrate the technical solutions of the embodiments of this specification, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some examples or embodiments of this specification. For those of ordinary skill in the art, without creative efforts, this specification can also be applied to other similar scenarios based on these drawings. Unless obvious from the language context or otherwise stated, the same reference numerals in the drawings represent the same structure or operation.

[0047] It should be understood that the "system", "device", "unit" and / or "module" used in this specification are a way to distinguish different components, elements, parts, portions or assemblies at different levels. However, if other words can achieve the same purpose, the said words can be replaced by other expressions.

[0048] As shown in this specification and the claims, unless the context clearly indicates an exception, words such as "a", "an", "one" and / or "the" are not specifically singular and may also include the plural. Generally speaking, the terms "including" and "comprising" only indicate the inclusion of the clearly identified steps and elements, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.

[0049] Flowcharts are used in this specification to illustrate the operations performed by the systems according to the embodiments of this specification. It should be understood that the operations before or after do not necessarily need to be executed precisely in sequence. On the contrary, they can be executed in reverse order or simultaneously. At the same time, other operations can also be added to these processes, or one or several operations can be removed from these processes.

[0050] The business risk identification method and system based on container shipping documents provided by the embodiments of this specification will be described in detail below with reference to the accompanying drawings.

[0051] Figure 1 It is an exemplary flowchart of a business risk identification method based on container shipping documents shown in some embodiments of this specification. In some embodiments, the business risk identification method based on container shipping documents can be executed by a processing logic, which can include hardware (e.g., circuits, dedicated logic, programmable logic, microcode, etc.), software (instructions running on a processing device to perform hardware emulation), etc., or any combination thereof. In some embodiments, Figure 1 One or more operations in the flowchart of the business risk identification method based on container shipping documents shown can be implemented by a processing device and / or a terminal device. For example, the business risk identification method based on container shipping documents can be stored in a storage device in the form of a computer program and / or instructions, and called and / or executed by a processing device and / or a terminal device. Referring to Figure 1 , in some embodiments, the business risk identification method based on container shipping documents can include:

[0052] Step S110, obtaining the document information to be verified related to the target container, where the document information to be verified is obtained by identifying at least one document related to the target container and arranging them in a preset order. In some embodiments, step S110 can be executed by the obtaining module 210 mentioned later.

[0053] A container is a large loading container with standardized dimensions and a sturdy structure. It is the core tool for international cargo transportation and can be used to transport various types of goods, thus ensuring the safety and transportation efficiency of the goods. In the actual sea transportation process, one container (each container has an irreplaceable container number) can correspond to one or more booking numbers, or one booking number can correspond to one or more containers, where the booking number can be understood as a business number.

[0054] In the embodiments of this application, the document information to be verified related to the target container can be obtained. The target container can be understood as a specific container, and its unique identifier is the container number. By obtaining all the document information related to this target container, the system can comprehensively master all the business data of this container during the sea transportation process. These document information include, but are not limited to, one or more of a booking form, a packing list, a bill of lading, a cargo list, an endorsement list, a commercial invoice, an insurance policy, a customs declaration form, a certificate of origin, an export license, a transportation contract, etc., covering the whole process from cargo packing, transportation to the destination. By identifying and extracting key information from these document information, a solid data foundation can be provided for subsequent risk identification and analysis.

[0055] Specifically, a booking note refers to a cargo consignment document prepared by a shipping company or its agent according to the booking requirements of the shipper. The packing list details information such as the types, quantities, and packing methods of the goods loaded into the container. The bill of lading is issued by the carrier or its agent, certifying that the goods have been received and promising to transport them to the designated location according to the contract terms. The cargo manifest, endorsement list, commercial invoice, insurance policy, etc. record the detailed information of the goods, transaction status, insurance status, etc. from different perspectives. The customs declaration form and certificate of origin are important bases for customs clearance and tax incentives. The export license ensures the legality of the exported goods. The transportation contract clarifies the rights and obligations between the carrier and the shipper, providing legal protection for both parties. There are usually interrelated information in these document information, jointly constituting the overall picture of the container shipping business. For example, the cargo information in the booking note should be consistent with the cargo information in the packing list to ensure that there are no omissions or errors during the packing process; the cargo information, quantity, destination, etc. in the bill of lading should match the information in the booking note and packing list to ensure that the carrier can transport the goods to the designated location according to the correct transportation requirements; the information such as the cargo price and quantity in the commercial invoice is closely related to the cargo value and quantity in the customs declaration form, directly affecting the customs clearance and tax calculation of the customs. Therefore, a comprehensive and accurate identification and correlation analysis of these document information is the key to identifying the risks of the container shipping business. By comprehensively obtaining and deeply analyzing this information, potential business risks can be discovered in a timely manner, such as inconsistent document information, lost goods, counterfeit and shoddy products mixed in, tax evasion, illegal transportation, etc. These risks may not only cause economic losses but also have a negative impact on the reputation and long-term development of the enterprise. Therefore, in the embodiments of this application, it is hoped to identify the anomalies and contradictions in the document information through intelligent algorithms, so as to accurately judge the existence of business risks.

[0056] Specifically, in some embodiments of this application, data of each document can be collected by scanning and then stored in a storage device. Further, the acquisition module 210 can obtain at least one document related to the target container from the storage device, then identify the at least one document and arrange them in a preset order, so as to obtain the document information to be verified related to the target container.

[0057] Specifically, in some embodiments, for each target document related to the target container (the target document can be any one of one or more documents related to the target container), the document type information of the target document, as well as the status information and / or content information of at least one sub-item included in the target document, can be extracted through OCR (Optical Character Recognition) technology to obtain the partial information to be verified corresponding to the target document. Herein, the sub-item can be understood as a specific field or entry in the target document, such as goods description, weight, volume, port of departure, port of destination, voyage number, shipper, carrier, consignee, notify party, freight payment method, etc. The partial information to be verified can be understood as partial data in the information of the document to be verified related to the target container. In the embodiments of the present application, the partial information to be verified can refer to the information set of all sub-items included in a target document.

[0058] In the embodiments of the present application, some sub-items in the target document can be used to represent the business status. For example, "confirmation of characteristic relationship", "confirmation of price impact relationship", and "confirmation of payment of royalties related to goods" that may be involved in the customs declaration form, etc. These business statuses can be reflected in the form of text (such as yes or no) or status codes (such as 1 or 0), and these texts or status codes can intuitively reflect the current status of the sub-item in the business process.

[0059] It should be noted that, in some embodiments of the present application, the various sub-items in the document can be divided by the tables in each document, or the aggregation, logic, and continuity between the texts. For example, different sub-items can be defined according to the column headers or row contents in the document table. For example, "goods description" may correspond to a certain column in the table, while "weight" and "volume" may respectively correspond to adjacent columns. At the same time, by analyzing the keywords or phrases in the text description, such as "port of departure", "port of destination", etc., they can also be accurately identified as independent sub-items. In addition, according to the business logic and the relevance between the documents, the sub-items in the document can be further subdivided and classified to improve the accuracy and efficiency of information recognition.

[0060] In some embodiments of the present application, a dedicated OCR recognition model for information recognition and extraction of various types of documents can be trained through machine learning. For example, in some embodiments, the initial OCR recognition model is trained with a large number of document samples to improve the accuracy and robustness of text and format recognition in container shipping documents. During the training process, the model will learn and understand the layout, text features, and possible logical relationships of different sub-items in the documents, so that it can more accurately identify and extract key information in actual applications. The specially trained OCR recognition model can greatly improve the accuracy and reliability of document information recognition and feature extraction, providing a reliable data basis for subsequent business risk assessment and decision-making. More training details about this OCR recognition model can be regarded as the prior art and will not be elaborated in detail in this specification.

[0061] Further, in some embodiments of the present application, the status information and / or content information of at least one sub-item included in each of the target documents can be arranged in a first preset order, and when the number of the documents is multiple, the partial information to be verified corresponding to the multiple documents can be arranged in a second preset order to obtain the information of the documents to be verified related to the target container.

[0062] Exemplarily, in some embodiments of the present application, the aforementioned first preset order may refer to sorting the data in each area according to the document layout or arrangement, and sorting each sub-item in each target document in the order from left to right and from top to bottom.

[0063] In some embodiments of the present application, the aforementioned second preset order can be obtained by sorting according to the importance of each document. Only as an example, in some embodiments, the main documents related to container shipping may include booking notes, bills of lading, packing lists, cargo manifests, endorsement lists, commercial invoices, insurance policies, customs declarations, certificates of origin, export licenses, etc. The second preset order can be expressed as: bill of lading > packing list > cargo manifest > endorsement list > commercial invoice > insurance policy > customs declaration > export license > certificate of origin > booking note. Through this sorting method, it can be ensured that during the business risk identification process, those documents that have a greater impact on the business process are given priority attention, thereby improving the efficiency and accuracy of risk identification. In addition, this ordered arrangement also helps relevant staff quickly locate key information during the manual verification process, reducing the time cost of searching for the required data among a large number of documents. It should be noted that in actual applications, the sorting rules of the documents can also be flexibly adjusted according to specific business scenarios and requirements to adapt to different business processes and risk management strategies.

[0064] In an embodiment of the present application, after determining the above-mentioned second preset order, corresponding partial information to be verified can be written into the corresponding position in the second preset order according to the document type information of at least one document related to the target container, and the missing data can be filled with a preset value to obtain the information of the documents to be verified related to the target container.

[0065] For example, in some embodiments, the documents related to the target container include a bill of lading, a packing list, a customs declaration form, and a commercial invoice in the actual business processing. Then, the information of the documents to be verified related to the target container can be expressed as: "data1-1, data1-2,..., data1-m; data2-1, data2-2,..., data2-n; null; null; data5-1, data5-2,..., data5-i; null; data7-1, data7-2,..., data7-j; null; null; null". Among them, "data1-1, data1-2,..., data1-m" represents the partial information to be verified obtained by sorting each sub-item in the bill of lading according to the aforementioned first preset order, where m is the total number of sub-items in the bill of lading; "data2-1, data2-2,..., data2-n" represents the partial information to be verified obtained by sorting each sub-item in the packing list according to the aforementioned first preset order, where n is the total number of sub-items in the packing list; "data5-1, data5-2,..., data5-i" represents the partial information to be verified obtained by sorting each sub-item in the commercial invoice according to the aforementioned first preset order, where i is the total number of sub-items in the commercial invoice; "data7-1, data7-2,..., data7-j" represents the partial information to be verified obtained by sorting each sub-item in the customs declaration form according to the aforementioned first preset order, where j is the total number of sub-items in the customs declaration form; null represents the result obtained by filling the missing data with a preset value. In other words, it represents the lack of document information at the corresponding position.

[0066] It should be noted that in the embodiments of the present application, the documents required for different business scenarios (such as different ports of destination, different types of goods, etc.) may be different. Some business scenarios may require other documents in addition to the types of documents listed above. In this case, the corresponding partial information to be verified can be arranged at the end of the second preset order, and the corresponding information of the documents to be verified can be obtained by arranging them in a similar manner. In some embodiments, at least one of the documents corresponding to the target container may include documents with different business numbers. In this case, the same type of documents can be arranged side by side in the second preset order, and the information of the documents to be verified related to the target container can be obtained based on this. In addition, it should also be noted that for some business scenarios, only some of the types of documents listed above may be required. In this case, the absence of other document information may not necessarily mean that there is a business risk.

[0067] Step S120: Perform feature extraction processing on the status information and / or content information of each sub-item in the information of the document to be verified to obtain an initial feature vector. In some embodiments, step S120 can be executed by the initial feature vector determination module 220 mentioned later.

[0068] In the embodiments of the present application, semantic encoding can be performed on the status information and / or content information of each sub-item in the information of the document to be verified processed by the above process through natural language processing to obtain at least one semantic vector corresponding to the information of the document to be verified, and the initial feature vector can be obtained based on the at least one semantic vector.

[0069] Specifically, in the embodiments of the present application, through natural language processing (Natural Language Processing, NLP), the text content in the information of the document to be verified can be converted into a numerical form that can be understood by a machine. For example, for the text information describing the status of goods, it can be converted into a numerical index indicating the quality of the goods status; for the text information describing the type of goods, it can be converted into a numerical vector representing the characteristics of the goods category.

[0070] In some embodiments of the present application, each sub-item in the above-mentioned document information to be verified can be converted into a corresponding semantic vector. In some other embodiments of the present application, multiple sub-items can also be converted into a semantic vector. Specifically, in some embodiments, the text information in each sub-item can be mapped to a high-dimensional vector space by means of word embeddings, so as to obtain its corresponding semantic vector. Exemplary word embedding methods can include Word2Vec, GloVe, BERT, etc. These word embedding methods can learn the semantic relationships between words from a large amount of text data, so as to map similar words to close positions in the vector space. In this way, the text content included in each sub-item in the document information to be verified can be converted into a numerical form that can be understood by a machine, so as to provide data support for subsequent business risk identification.

[0071] In the embodiments of the present application, the semantic vectors corresponding to each sub-item can be arranged in the above-mentioned first preset order and second preset order, and then the above-mentioned initial feature vector can be obtained. In the embodiments of the present application, this initial feature vector can be understood as a vector represented in numerical form that includes the status information and / or content information of each sub-item in the document information to be verified. This initial feature vector will be used as important input data in the subsequent business risk identification step, and business risk prediction and assessment will be carried out based on this.

[0072] Step S130, perform deep feature mining at multiple different levels on the initial feature vector to obtain multiple deep feature vectors corresponding to the document information to be verified, where each deep feature vector is used to represent the deep feature of the document information to be verified at the corresponding level. In some embodiments, step S130 can be executed by the deep feature vector determination module 230 mentioned later.

[0073] It should be noted that for the same or related content in different documents, it should correspond to the same or similar semantic vector. Therefore, in this initial feature vector, it will be represented in the same or similar numerical form. However, only by identifying the same or related information in this initial feature vector, it is still impossible to confirm in what circumstances there is a business risk in this initial feature vector and in what circumstances there is no business risk. Based on this, in some embodiments of the present application, deep feature mining can be performed on the initial feature vector at multiple different levels to obtain multiple deep feature vectors corresponding to the document information to be verified, where each deep feature vector is used to represent the deep feature of the document information to be verified at the corresponding level.

[0074] It can be understood that by performing multiple different levels of deep feature mining on the initial feature vector, more abstract and higher-level feature representations can be gradually extracted, thereby capturing more complex feature patterns and semantic information. Specifically, in some embodiments of the present application, the initial feature vector can be input into a deep feature extraction network, wherein the deep feature extraction network includes multiple hidden layers, each hidden layer corresponds to an activation function, and each of the hidden layers and the activation function is used to perform one level of deep feature mining on the initial feature vector; further, based on the outputs of the multiple hidden layers, multiple deep feature vectors corresponding to the document information to be verified can be obtained.

[0075] Figure 2 is a schematic diagram of a feature vector processing flow according to some embodiments of this specification. Figure 2 In the embodiment of the present application, each hidden layer of the deep feature extraction network (e.g. Figure 2 H1, H2, H3, ..., Hx) shown may include multiple neurons (the neuron can be understood as a computing unit, and multiple neurons can constitute a convolution kernel. In the embodiment of the present application, different hidden layers can have the same or different numbers of neurons). Each neuron can receive the output from the neurons in the previous layer, and calculate based on the weights and biases to output a result. These results can be passed to the next layer of neurons after nonlinear transformation by the activation function (such as ReLU), so as to perform higher-level deep feature extraction. In this way, the deep feature extraction network can gradually extract the complex features in the document information to be verified, which helps to dig out deep data information, thereby realizing accurate identification of business risks.

[0076] Specifically, in the embodiment of the present application, the first hidden layer H1 can accept the initial feature vector and process the initial feature vector through the multiple neurons it contains to obtain the output of each neuron. Furthermore, the output of each neuron in the first hidden layer H1 can be used as the input of the second hidden layer H2, and each neuron in the second hidden layer H2 processes these inputs, and so on, until the last hidden layer Hx (x is the number of hidden layers). In each hidden layer, by continuously adjusting the weights and biases of the neurons, the deep feature extraction network can learn more abstract and complex feature representations.

[0077] In some embodiments of the present application, for each hidden layer, the outputs of multiple neurons included therein can be integrated to obtain the output corresponding to this hidden layer. For example, by performing operations such as weighted averaging or summing the outputs of each neuron in the first hidden layer H1, the depth feature vector O1 corresponding to the first hidden layer H1 can be obtained. Similarly, for subsequent hidden layers H2, H3... Hx, the same method can be used to integrate the outputs of multiple neurons included in each of them to obtain the corresponding depth feature vectors O2, O3... Ox. It can be understood that these depth feature vectors can more comprehensively reflect the complex features in the document information to be verified, thereby providing strong support for the accurate identification of subsequent business risks.

[0078] Step S140: Perform feature fusion on the multiple depth feature vectors corresponding to the document information to be verified to obtain a fused feature vector. In some embodiments, step S140 can be executed by the fused feature vector determination module 240 mentioned later.

[0079] Continuing to refer to Figure 2 , after obtaining the above-mentioned depth feature vectors O1, O2, O3... Ox, the multiple depth feature vectors can be subjected to feature fusion by a feature fusion unit to obtain a fused feature vector.

[0080] Specifically, in some embodiments of the present application, the fused feature vector can be obtained by concatenating the above-mentioned multiple depth feature vectors O1, O2, O3... Ox.

[0081] It can be understood that in the embodiments of the present application, this fused feature vector contains the depth features obtained by performing feature extraction at different levels on the document information to be verified. Therefore, it can more comprehensively express the complex features in the document information to be verified. Based on this, in the subsequent business risk identification process, more accurate and comprehensive analysis can be performed based on this fused feature vector, thereby improving the accuracy and reliability of business risk identification. In some embodiments of the present application, the feature fusion unit can be implemented by a dedicated hardware circuit or a software module to achieve efficient fusion of multiple depth feature vectors.

[0082] Step S150: Used to perform risk identification in multiple aspects based on the fused feature vector to obtain corresponding business risk identification results. In some embodiments, step S150 can be executed by the business risk identification module 250 mentioned later.

[0083] In the embodiments of the present application, the fused feature vector obtained in the foregoing process can be respectively input into multiple pre-trained classifiers for classification processing, and the business risk identification results can be obtained based on the outputs of the multiple classifiers. Among them, each of the classifiers corresponds to a risk identification in one aspect.

[0084] Specifically, since the risks reflected by container shipping documents may cover multiple aspects, such as compliance risks (whether in line with relevant laws and regulations of international and domestic trade and transportation), financial risks (involving financial verification issues such as invoices and payments), transportation risks (possible delays, damages or losses in links such as goods transportation, unloading and customs clearance), contract risks (discrepancies between goods shipment and delivery terms and the contract, or performance issues of the transportation party), and information accuracy risks (for example, whether the information in the documents is accurate, such as whether the product name, quantity, transportation mode, etc. are consistent with the actual situation).

[0085] In the embodiments of the present application, to ensure the accuracy and reliability of risk identification, multiple classifiers can be used to process the fused feature vector respectively. Among them, each classifier is specially trained (for example, trained specifically using the document data of relevant cases), so that each classifier can identify specific types of risks. It should be noted that in the embodiments of the present application, by using relevant cases to train each classifier differently, each classifier can have different focuses during data processing. For example, the classifier for compliance risks may focus on whether the documents comply with relevant laws and regulations of international trade and transportation after training; the classifier for financial risks may pay more attention to the verification of financial information such as invoices and payments after training; the classifier for transportation risks will evaluate possible problems in links such as goods transportation, unloading and customs clearance; the classifier for contract risks may focus on analyzing whether the goods shipment and delivery terms are consistent with the contract and the performance of the transportation party; while the classifier for information accuracy risks may focus on whether the information in the documents such as product name, quantity, transportation mode, etc. is consistent with the actual situation.

[0086] It should be pointed out that in the embodiments of the present application, through the collaborative work of multiple classifiers, comprehensive and accurate identification of various risks involved in container shipping documents can be achieved. This design of multiple classifiers not only improves the efficiency of risk identification but also ensures the accuracy of identification. In addition, since each classifier is specially trained and has different focuses, it can better adapt to different types of risk identification requirements, making the entire risk identification system more flexible and efficient.

[0087] Figure 3 is a schematic diagram of the business risk identification principle shown in some embodiments of this specification. Refer to Figure 3, in some embodiments of the present application, the fused feature vector can be input into classifier 1, classifier 2, …, classifier y respectively for processing to obtain the classification results F1 output by classifier 1, the classification results F2 output by classifier 2, …, the classification results Fy output by classifier y. In some embodiments of the present application, classifier 1, classifier 2, …, classifier y can be binary classifiers, and their output results are 0 or 1, where 0 indicates that the document does not involve the risk identified by the corresponding classifier, and 1 indicates that the document involves the risk identified by the corresponding classifier. In some embodiments of the present application, classifier 1, classifier 2, …, classifier y can also be linear classifiers (such as softmax classifiers), and their output results can be represented in the form of probabilities. For example, the output result of each classifier can be in the range of [0, 1], where 0 indicates that the information of the document to be verified does not involve the risk identified by the corresponding classifier, 0.1 indicates that the probability that the information of the document to be verified involves the risk identified by the corresponding classifier is small, and 0.9 indicates that the probability that the information of the document to be verified involves the risk identified by the corresponding classifier is large.

[0088] The training methods of each classifier involved in the embodiments of the present application are briefly introduced below:

[0089] In the embodiments of the present application, a sample fused feature vector obtained by processing at least one sample document related to a sample case can be obtained, as well as a plurality of risk labels corresponding to the sample fused feature vector, where each risk label in the plurality of risk labels is used as the expected output of a classifier (the risk label can be obtained through manual evaluation); then, the sample fused feature vector can be input into the plurality of classifiers respectively to obtain the predicted output corresponding to each classifier; further, based on the difference between the predicted output and the expected output corresponding to each classifier, the loss value of each classifier can be determined; finally, based on the loss value, the plurality of classifiers are iteratively optimized respectively until the trained plurality of classifiers are obtained when the preset training conditions are met.

[0090] Exemplarily, in some embodiments, the loss value can be expressed as:

[0091]

[0092] where L represents the sum of the loss values of all classifiers, represents the predicted output of the i-th classifier, represents the corresponding risk label, and y represents the total number of classifiers.

[0093] In some embodiments of the present application, in order to accelerate the convergence speed of classifier training to a certain extent, each classifier can be trained separately. For example, for the classifier for compliance risk identification, a large number of cases related to compliance risks can be obtained (including positive samples and negative samples, that is, the sample cases need to include both documentary data with risks and documentary data without risks). These cases can include one or more sample documents related to the sample container. In the embodiments of the present application, the one or more sample documents related to the sample container can be processed in the above manner to obtain the corresponding sample fusion feature vector and the compliance risk label corresponding to the sample fusion feature vector. Then, the sample fusion feature vector can be input into the corresponding target classifier respectively to obtain the prediction output corresponding to the target classifier. Further, based on the difference between the prediction output corresponding to the target classifier and the expected output, its loss value can be determined. Finally, based on the loss value, the target classifier is iteratively optimized until the training of the target classifier is completed when the preset training conditions are met.

[0094] In the embodiments of the present application, the gradient descent method or other optimization algorithms can be used to iteratively optimize the classifier. Specifically, in each iteration, the loss value is calculated according to the parameters of the current classifier, and then the gradient of the loss value with respect to the classifier parameters is calculated through the backpropagation algorithm. Then, the classifier parameters are updated according to the gradient to reduce the loss value. By continuously repeating this process until the loss value converges to a small value or reaches the preset number of iterations, it can be considered that the classifier has been trained. It can be understood that after training, each classifier can obtain the corresponding generalization ability and can accurately predict the business risks of the documentary information to be verified related to the target container. Moreover, through this method, trained classifiers for different types of risks (such as compliance risks, financial risks, contract risks, etc.) can be obtained, and these classifiers can be used to classify the fusion feature vectors processed in the above process to accurately identify the potential business risks in the container shipping documents.

[0095] In some embodiments of the present application, the trained classifier can be used to process the above fusion feature vector to obtain the corresponding business risk identification result. The business risk identification result can include multiple data (for example, y data), and each data is used to reflect the risk situation in one aspect. In some embodiments of the present application, when the classification result obtained by using the trained classifier to process the above fusion feature vector reflects the existence of at least one risk, the relevant data can be backed up and recorded and sent to the relevant business personnel for manual review.

[0096] Specifically, in some embodiments of the present application, the system can automatically mark the documents with risks and save them together with the relevant business data to a specified storage medium for subsequent reference and analysis. At the same time, in some embodiments, the system can also generate risk warning notifications and send them to the corresponding business personnel in a timely manner via email, text message, or system message. After receiving the warning notification, the business personnel can log in to the system to view the detailed risk information and relevant documents for manual review. During the review process, the business personnel can further confirm and evaluate the risks identified by the system by combining their professional knowledge and experience to ensure the accuracy and reliability of risk identification. It should be noted that in this way, the efficiency and accuracy of risk identification in the container shipping document business can be effectively improved, while reducing the operational risks of the enterprise and greatly reducing the labor costs required for document review.

[0097] Figure 4 is a schematic diagram of the business risk identification principle according to some other embodiments of this specification. Referring to Figure 4 In some embodiments of the present application, for each of the classifiers, the attention mechanism can be used to perform attention encoding on the sample fusion feature vector to obtain an attention fusion vector corresponding to each classifier; then, the attention fusion vector is used as the input data of the classifier and input into the corresponding classifier for processing to obtain a prediction output corresponding to each classifier.

[0098] It can be understood that by performing attention encoding on the sample fusion feature vector through the attention mechanism, the classifier can pay more attention to the features that have a greater impact on the risk identification result, thereby improving the accuracy and efficiency of risk identification. Specifically, the attention mechanism can assign different weights to different features according to their importance, so that important features can receive more attention and consideration during the risk identification process. In this way, the risk points reflected by the different-level deep features in the fusion feature vector corresponding to the document information to be verified can be captured more accurately, thereby improving the efficiency and accuracy of risk identification.

[0099] At the same time, since the attention mechanism can adaptively adjust the feature weights according to different situations, it can also improve the flexibility and adaptability of risk identification. More details about performing attention encoding on the sample fusion feature vector through the attention mechanism can be regarded as prior art and will not be elaborated in this specification.

[0100] Figure 5 is a schematic diagram of the modules of the business risk identification system based on container shipping documents according to some embodiments of this specification. In some embodiments, Figure 5The business risk identification system 200 based on container shipping documents shown can be implemented in the form of software and / or hardware. For example, it can be configured into a processing device and / or a terminal device in the form of software and / or hardware to process at least one document related to a target container and identify potential business risks based on the at least one document.

[0101] Referring to Figure 5 , in some embodiments, the business risk identification system 200 based on container shipping documents may include an acquisition module 210, an initial feature vector determination module 220, a deep feature vector determination module 230, a fused feature vector determination module 240, and a business risk identification module 250.

[0102] The acquisition module 210 can be used to acquire the information of the documents to be verified related to the target container, and the information of the documents to be verified is obtained by identifying at least one document related to the target container and arranging them in a preset order.

[0103] The initial feature vector determination module 220 can be used to perform feature extraction processing on the status information and / or content information of each sub-item in the information of the documents to be verified to obtain an initial feature vector.

[0104] The deep feature vector determination module 230 can be used to perform deep feature mining at multiple different levels on the initial feature vector to obtain multiple deep feature vectors corresponding to the information of the documents to be verified, where each deep feature vector is used to represent the deep feature of the information of the documents to be verified at the corresponding level.

[0105] The fused feature vector determination module 240 can be used to perform feature fusion on multiple deep feature vectors corresponding to the information of the documents to be verified to obtain a fused feature vector.

[0106] The business risk identification module 250 can be used to perform risk identification in multiple aspects based on the fused feature vector to obtain corresponding business risk identification results.

[0107] For more details about the above-mentioned various modules, reference can be made to other parts of this specification (such as Figures 1 to 4 the relevant descriptions in that part), which will not be elaborated here.

[0108] It should be understood that Figure 5The business risk identification system 200 based on container shipping documents shown and its modules can be implemented in various ways. For example, in some embodiments, the system and its modules can be implemented through hardware, software, or a combination of software and hardware. Among them, the hardware part can be implemented using dedicated logic; the software part can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. Those skilled in the art can understand that the above methods and systems can be implemented using computer-executable instructions and / or included in processor control code. For example, such code is provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The systems and modules in this specification can be implemented not only by hardware circuits such as very large scale integrated circuits or gate arrays, semiconductors such as logic chips and transistors, or programmable hardware devices such as field programmable gate arrays and programmable logic devices, but also by software executed by various types of processors, or by a combination of the above hardware circuits and software (e.g., firmware).

[0109] It should be noted that the above description of the business risk identification system 200 based on container shipping documents is provided for illustrative purposes only and is not intended to limit the scope of this specification. It can be understood that for those skilled in the art, according to the description of this specification, without departing from this principle, various modules can be arbitrarily combined, or a subsystem can be formed and connected to other modules. For example, Figure 5 the acquisition module 210, the initial feature vector determination module 220, the deep feature vector determination module 230, the fused feature vector determination module 240, and the business risk identification module 250 described in can be different modules in a system, or a single module can implement the functions of two or more of the above modules. Such variations are all within the protection scope of this specification.

[0110] In summary, the beneficial effects that the embodiments of this specification may bring include, but are not limited to: (1) In the business risk identification method and system based on container shipping documents provided in some embodiments of this specification, by extracting and processing the characteristics of the documents to be verified to obtain an initial feature vector, and then performing deep feature mining at multiple different levels on this initial feature vector to obtain multiple deep feature vectors, it can more comprehensively reflect the complex characteristics in the documents to be verified, thereby improving the accuracy and reliability of the business risk identification results obtained based on these multiple deep feature vectors in the subsequent process; (2) In the business risk identification method and system based on container shipping documents provided in some embodiments of this specification, by using multiple classifiers to process the fused feature vector obtained by fusing multiple deep feature vectors at different levels, realizing risk identification in multiple aspects, it can enable each classifier to have different focuses of attention after training, thus better adapting to different types of risk identification requirements, improving the efficiency of risk identification and the accuracy of risk identification results to a certain extent, and making the entire risk identification system more flexible and efficient; (3) In the business risk identification method and system based on container shipping documents provided in some embodiments of this specification, by implementing container shipping document verification through artificial intelligence, it can effectively improve the efficiency and accuracy of business risk identification of container shipping documents, while reducing the operational risk of enterprises and greatly reducing the labor cost required for document review.

[0111] It should be noted that the beneficial effects that different embodiments may bring are different. In different embodiments, the beneficial effects that may be produced can be any one or several combinations of the above, or any other beneficial effects that may be obtained.

[0112] The basic concepts have been described above. Obviously, for those skilled in the art, the above detailed disclosure is only an example and does not constitute a limitation to this specification. Although not explicitly stated here, those skilled in the art may make various modifications, improvements, and corrections to this specification. Such modifications, improvements, and corrections are proposed in this specification, so such modifications, improvements, and corrections still fall within the spirit and scope of the exemplary embodiments of this specification.

[0113] At the same time, this specification uses specific terms to describe the embodiments of this specification. Such as "one embodiment", "an embodiment", and / or "some embodiments" mean a certain feature, structure, or characteristic related to at least one embodiment of this specification. Therefore, it should be emphasized and noted that "an embodiment" or "one embodiment" or "an alternative embodiment" mentioned twice or more at different positions in this specification is not necessarily referring to the same embodiment. In addition, certain features, structures, or characteristics in one or more embodiments of this specification can be appropriately combined.

[0114] In addition, those skilled in the art can understand that various aspects of this specification can be illustrated and described by several patentable types or situations, including any new and useful process, machine, product, or composition of matter, or any new and useful improvement thereof. Accordingly, various aspects of this specification can be executed entirely by hardware, entirely by software (including firmware, resident software, microcode, etc.), or by a combination of hardware and software. The above hardware or software can all be referred to as "data blocks", "modules", "engines", "units", "components", or "systems". In addition, various aspects of this specification may be embodied as a computer product located in one or more computer-readable media, which includes computer-readable program code.

[0115] A computer storage medium may contain a propagated data signal having computer program code therein, for example, on a baseband or as part of a carrier wave. This propagated signal may have various manifestations, including electromagnetic form, optical form, etc., or a suitable combination thereof. A computer storage medium can be any computer-readable medium other than a computer-readable storage medium, which can be connected to an instruction execution system, apparatus, or device to effect communication, propagation, or transmission for use of a program. The program code located on a computer storage medium can be propagated through any suitable medium, including radio, cable, fiber optic cable, RF, or similar media, or any combination of the above media.

[0116] The computer program code required for the operation of each part of this specification can be written in any one or more programming languages, including object-oriented programming languages such as Java, Scala, Smalltalk, Eiffel, JADE, Emerald, C++, C#, VB.NET, Python, etc., conventional procedural programming languages such as C language, Visual Basic, Fortran2003, Perl, COBOL2002, PHP, ABAP, dynamic programming languages such as Python, Ruby, and Groovy, or other programming languages. The program code can run entirely on the user's computer, or as an independent software package on the user's computer, or partly on the user's computer and partly on a remote computer, or entirely on a remote computer or processing device. In the latter case, the remote computer can be connected to the user's computer through any network form, such as a local area network (LAN) or a wide area network (WAN), or connected to an external computer (for example, through the Internet), or in a cloud computing environment, or used as a service such as software as a service (SaaS).

[0117] In addition, unless explicitly stated in the claims, the order of the processing elements and sequences, the use of numerical and alphabetical characters, or the use of other names described in this specification are not used to limit the order of the processes and methods in this specification. Although some currently useful embodiments of the invention are discussed through various examples in the above disclosure, it should be understood that such details are for illustrative purposes only. The appended claims are not limited to the disclosed embodiments. On the contrary, the claims are intended to cover all modifications and equivalent combinations that conform to the essence and scope of the embodiments of this specification. For example, although the system components described above can be implemented by hardware devices, they can also be implemented only through software solutions, such as installing the described system on existing processing devices or mobile devices.

[0118] Similarly, it should be noted that, in order to simplify the presentation of the disclosure in this specification and thus help the understanding of one or more embodiments of the invention, in the foregoing description of the embodiments of this specification, multiple features are sometimes grouped into one embodiment, drawing, or description thereof. However, this method of disclosure does not mean that the features required by the subject matter of this specification are more than those mentioned in the claims. In fact, the features of the embodiments are fewer than all the features of the individual embodiments disclosed above.

[0119] In some embodiments, numbers are used to describe components and the quantity of attributes. It should be understood that such numbers used to describe the embodiments are modified by the modifiers "about", "approximate", or "substantially" in some examples. Unless otherwise stated, "about", "approximate", or "substantially" indicate that the stated number allows a variation of ±20%. Accordingly, in some embodiments, the numerical parameters used in the specification and claims are approximate values, and such approximate values may vary according to the characteristics required by individual embodiments. In some embodiments, the numerical parameters should consider the specified significant digits and adopt the method of retaining the general number of digits. Although the numerical ranges and parameters used in some embodiments of this specification to confirm the breadth of their scope are approximate values, in specific embodiments, such numerical settings are made as precise as possible within the feasible range.

[0120] For each patent, patent application, patent application publication, and other materials cited in this specification, such as articles, books, specifications, publications, documents, etc., their entire contents are hereby incorporated into this specification by reference. This excludes the application history documents that are inconsistent with or conflict with the content of this specification, and also excludes the documents that limit the broadest scope of the claims of this specification (currently or subsequently appended to this specification). It should be noted that if there are inconsistencies or conflicts between the descriptions, definitions, and / or uses of terms in the supplementary materials of this specification and the content described in this specification, the descriptions, definitions, and / or uses of terms in this specification shall prevail.

[0121] Finally, it should be understood that the embodiments described in this specification are only used to illustrate the principles of the embodiments of this specification. Other variations may also fall within the scope of this specification. Therefore, by way of example and not limitation, alternative configurations of the embodiments of this specification may be regarded as consistent with the teachings of this specification. Accordingly, the embodiments of this specification are not limited to the embodiments explicitly presented and described in this specification.

Claims

1. A business risk identification method based on container shipping documents, characterized in that: The method comprises: Acquire document information to be checked related to the target container, wherein the document information to be checked is obtained by identifying at least one document related to the target container and arranging it in a preset order, wherein the at least one document includes one or more of a booking list, a packing list, a bill of lading, a cargo list, an annotated list, a commercial invoice, an insurance policy, a customs declaration, a certificate of origin, an export license, and a transport contract; Performing feature extraction processing on the status information and / or content information of each sub-item in the document information to be verified to obtain an initial feature vector; Performing deep feature mining at multiple different levels on the initial feature vector to obtain multiple deep feature vectors corresponding to the document information to be verified, wherein each of the deep feature vectors is used to represent the deep features of the document information to be verified at a corresponding level; Performing feature fusion on multiple deep feature vectors corresponding to the document information to be verified to obtain a fused feature vector; Based on the fused feature vector, multiple aspects of risk identification are performed to obtain corresponding business risk identification results; The risks in the multiple aspects include compliance risk, financial risk, transportation risk, contract risk and information accuracy risk; the risk identification in the multiple aspects based on the fused feature vector to obtain the corresponding business risk identification result includes: inputting the fused feature vector into a plurality of pre-trained classifiers for classification processing, and obtaining the business risk identification result based on the outputs of the plurality of classifiers; wherein each of the classifiers corresponds to risk identification in one aspect; the plurality of classifiers are trained based on the following method: Acquire a sample fusion feature vector obtained by processing at least one sample document related to the sample case, and a plurality of risk labels corresponding to the sample fusion feature vector, wherein each of the plurality of risk labels is used as an expected output of a classifier; For each of the classifiers, the sample fusion feature vector is attention-encoded by an attention mechanism to obtain an attention fusion vector corresponding to each classifier; Input the attention fusion vector as the input data of the classifier into the corresponding classifier for processing, and obtain the prediction output corresponding to each classifier; Determine a loss value for each classifier based on a difference between a predicted output and an expected output corresponding to each of the classifiers; The multiple classifiers are iteratively optimized based on the loss values ​​until the multiple trained classifiers are obtained when preset training conditions are met.

2. The business risk identification method based on container shipping documents according to claim 1 is characterized in that: The obtaining of the document information to be checked related to the target container includes: For each target document related to the target container, extract the document type information of the target document and the status information and / or content information of at least one sub-item contained in the target document by using OCR recognition technology to obtain the partial information to be verified corresponding to the target document; The status information and / or content information of at least one sub-item contained in each of the target documents is arranged according to a first preset order, and when there are multiple documents, the local information to be verified corresponding to the multiple documents are arranged according to a second preset order to obtain the document information to be verified related to the target container.

3. The business risk identification method based on container shipping documents according to claim 2 is characterized in that: The step of arranging the partial information to be checked corresponding to the plurality of documents in accordance with a second preset order to obtain the document information to be checked related to the target container includes: Sorting the documents according to their importance to determine the second preset order; According to the document type information of at least one document related to the target container, the corresponding local information to be verified is written into the corresponding position in the second preset order, and the missing data is filled with preset values ​​to obtain the document information to be verified related to the target container.

4. The business risk identification method based on container shipping documents according to claim 1 is characterized in that: The feature extraction process is performed on the status information and / or content information of each sub-item in the document information to be verified to obtain an initial feature vector, including: The status information and / or content information of each sub-item in the document information to be verified is semantically encoded through natural language processing to obtain at least one semantic vector corresponding to the document information to be verified, and the initial feature vector is obtained based on the at least one semantic vector.

5. The business risk identification method based on container shipping documents according to claim 1 is characterized in that: The performing of deep feature mining at multiple different levels on the initial feature vector to obtain multiple deep feature vectors corresponding to the document information to be verified includes: Inputting the initial feature vector into a deep feature extraction network, wherein the deep feature extraction network includes a plurality of hidden layers, each hidden layer corresponds to an activation function, and each of the hidden layers and the activation function is used to perform a level of deep feature mining on the initial feature vector; Based on the outputs of the multiple hidden layers, multiple deep feature vectors corresponding to the document information to be verified are obtained.

6. The business risk identification method based on container shipping documents according to claim 5 is characterized in that: The step of fusing the multiple deep feature vectors corresponding to the document information to be verified to obtain the fused feature vector includes: concatenating the multiple deep feature vectors to obtain the fused feature vector.

7. A business risk identification system based on container shipping documents, characterized in that: The system comprises: an acquisition module, used for acquiring document information to be checked related to the target container, wherein the document information to be checked is obtained by identifying at least one document related to the target container and arranging it in a preset order, wherein the at least one document includes one or more of a booking list, a packing list, a bill of lading, a cargo list, an annotated list, a commercial invoice, an insurance policy, a customs declaration, a certificate of origin, an export license, and a transport contract; An initial feature vector determination module is used to perform feature extraction processing on the status information and / or content information of each sub-item in the document information to be verified to obtain an initial feature vector; A deep feature vector determination module is used to perform deep feature mining at multiple different levels on the initial feature vector to obtain multiple deep feature vectors corresponding to the document information to be verified, wherein each of the deep feature vectors is used to represent the deep features of the document information to be verified at a corresponding level; A fused feature vector determination module is used to perform feature fusion on multiple deep feature vectors corresponding to the document information to be verified to obtain a fused feature vector; A business risk identification module, used to identify risks in multiple aspects based on the fused feature vector to obtain corresponding business risk identification results; The risks in the multiple aspects include compliance risk, financial risk, transportation risk, contract risk and information accuracy risk; the business risk identification module is specifically used to: input the fusion feature vectors into multiple pre-trained classifiers for classification processing, and obtain the business risk identification results based on the outputs of the multiple classifiers; each of the classifiers corresponds to one aspect of risk identification; the multiple classifiers are trained based on the following method: Acquire a sample fusion feature vector obtained by processing at least one sample document related to the sample case, and a plurality of risk labels corresponding to the sample fusion feature vector, wherein each of the plurality of risk labels is used as an expected output of a classifier; For each of the classifiers, the sample fusion feature vector is attention-encoded by an attention mechanism to obtain an attention fusion vector corresponding to each classifier; Input the attention fusion vector as the input data of the classifier into the corresponding classifier for processing, and obtain the prediction output corresponding to each classifier; Determine a loss value for each classifier based on a difference between a predicted output and an expected output corresponding to each of the classifiers; The multiple classifiers are iteratively optimized based on the loss values ​​until the multiple trained classifiers are obtained when preset training conditions are met.

Citation Information

Patent Citations

  • Business risk prediction method and device, computer equipment and storage medium

    CN114662570A

  • Method and device for realizing risk data identification processing in credential environment, processor and computer readable storage medium thereof

    CN116150658A