Risk detection method and apparatus, and storage medium
Through machine learning models, optimize the ingredient content information of edible products, combine it with food safety standards and access knowledge base, the problem of inaccurate manual detection is solved, objective risk assessment and cause analysis are realized, and detection efficiency and accuracy are improved.
Patent Information
- Application Number
- PCT/CN2024/142756
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-28
- Filing Date
- 2024-12-26
- Publication Date
- 2025-07-03
AI Technical Summary
In the prior art, the risk detection of imported edible commodities mainly relies on manual methods and is easily affected by human factors, resulting in inadequate detection results.
The machine learning model is used to identify the component content information in the customs clearance data, and the corresponding relationship between the ingredient name and content data is optimized through the BERT, BILSTM and CRF models, and combined with the food safety standard library and the access knowledge base, the first risk value, the second risk value and the third risk value are generated to finally determine the import risk value.
It realizes objective and accurate risk detection results, saves labor costs, improves detection efficiency, and provides information on risk causes to facilitate problem positioning.
Smart Images

Figure CN2024142756_03072025_PF_FP_ABST
Abstract
Description
Risk detection method and device, storage medium
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This disclosure is based on and claims priority to an application with CN application number 202311830704.7 and filing date December 28, 2023. The disclosure content of this CN application is hereby incorporated into this disclosure as a whole. Technical Field
[0003] The present disclosure relates to the field of security access, and in particular to a risk detection method and device, and a storage medium. Background Art
[0004] Safe access is a business scenario of great concern to customs, and is a common area of business alongside tax collection and management, bonded research and development, RCEP (Regional Comprehensive Economic Partnership), epidemic prevention and control, and cross-border e-commerce. Within this context, particularly during import security inspections, inspections of edible commodity declarations often raise numerous issues, including discrepancies or incomplete documentation, non-compliant origins, excessive ingredient content, illegal use of certain substances, and expired goods. Safe access supervision for edible goods also faces numerous challenges.
[0005] At present, in order to ensure import safety, import risk testing of edible goods is mainly carried out manually. Summary of the Invention
[0006] In a first aspect of the present disclosure, a risk detection method is provided, which is performed by a risk detection device and includes: identifying ingredient content information of edible commodities included in customs clearance data to obtain multiple ingredient names and corresponding content data; comparing the multiple ingredient names and corresponding content data with a predetermined food safety standard library to determine a first risk value of the edible commodity; determining a second risk value of multiple pieces of access information of the edible commodity included in the customs clearance data based on a food access knowledge base; obtaining historical risk information of the edible commodity; determining a third risk value of the edible commodity based on the historical risk information; and determining the import risk value of the edible commodity based on the first risk value, the second risk value and the third risk value.
[0007] In some embodiments, determining the import risk value of the edible commodity includes determining the import risk value of the edible commodity according to the sum of the first risk value, the second risk value, and the third risk value.
[0008] In some embodiments, identifying the ingredient content information of edible products included in customs clearance data includes: processing the ingredient content information using a first machine learning model to obtain embedded information; processing the embedded information using a second machine learning model to obtain a feature vector; processing the feature vector using a third machine learning model to obtain a recognition result, wherein the recognition result includes the multiple ingredient names and the multiple content data; and optimizing the correspondence between the multiple ingredient names and the multiple content data.
[0009] In some embodiments, optimizing the correspondence between the multiple component names and the multiple content data includes: adding punctuation marks after each component name and each content data in the recognition result; if the first information to be processed is one of the component name and content data, and the second information after the first information is the other of the component name and content data, then the punctuation mark after the second information is used as the grouping boundary, and the first information and the second information are corresponded.
[0010] In some embodiments, optimizing the correspondence between the multiple component names and the multiple content data includes: when there are brackets in the recognition result, and there are component names inside the brackets, and the two sides outside the brackets have component names and content data adjacent to the brackets, respectively, identifying the content inside the brackets and establishing a correspondence between the component names and content data adjacent to the brackets.
[0011] In some embodiments, optimizing the correspondence between the multiple component names and the multiple content data includes: generating first identification information and second identification information when the recognition result includes a connector, and the left end of the connector has first content data and the right end of the connector has second content data, wherein the first identification information includes information that the component name associated with the connector is greater than the first content data, and the second identification information includes information that the component name associated with the connector is less than the second content data.
[0012] In some embodiments, optimizing the correspondence between the multiple component names and the multiple content data includes: adding punctuation marks after each component name and each content data in the recognition result; using the punctuation marks as grouping boundaries to obtain multiple groups, wherein, if a group only includes one component name, then adding predetermined content data after the component name; if a group only includes one content data, then adding the predetermined component name before the content data; if in adjacent first and second groups, the first group includes the first component name and the predetermined content data, and the second group includes the predetermined component name and the second content data, then merging the first group and the second group, deleting the predetermined component name and the predetermined content data, so that a correspondence is established between the first component name and the second content data.
[0013] In some embodiments, the first machine learning model is a BERT model; the second machine learning model is a BILSTM model; and the third machine learning model is a CRF model.
[0014] In some embodiments, the access information includes whether the edible commodity has exceeded its shelf life, and at least one of the following: whether the country of origin of the edible commodity is allowed to enter, whether the edible commodity has a certificate of origin, whether the consumer use unit of the edible commodity has an import license, whether the manufacturer of the edible commodity is validly registered, whether the manufacturer of the edible commodity has a production and processing license, and whether the manufacturer of the edible commodity has an inspection and quarantine certificate.
[0015] In some embodiments, the historical risk information includes at least one of the following: the historical risk of the country of origin of the edible commodity, the historical risk of the manufacturer of the edible commodity, the historical risk of the edible commodity, the packaging discreteness of the edible commodity, the historical number of order changes of the declaring unit of the edible commodity, and the degree of expiration of the edible commodity.
[0016] In some embodiments, determining the third risk value of the edible commodity based on the historical risk information includes: processing the historical risk information using an integrated learning model to obtain the third risk value.
[0017] In some embodiments, a first risk vector is generated, wherein the i-th element in the first risk vector corresponds to the risk value of the i-th access information among the multiple access information, 1≤i≤N, and N is the total number of the multiple access information; the first risk value and the third risk value are inserted into the first risk vector to obtain a second risk vector.
[0018] In some embodiments, first risk description information of the edible commodity is generated, wherein the first risk description information includes ingredients whose content does not meet the requirements of the food safety standard library; second risk description information of the edible commodity is generated, wherein the second risk description information includes access information that does not meet the requirements of the food access knowledge base; and risk reason information of the edible commodity is generated based on the first risk description information and the second risk description information.
[0019] In a second aspect of the present disclosure, a risk detection device is provided, comprising: a memory; and a processor coupled to the memory, the processor being configured to execute the method as described in any of the above embodiments based on instructions stored in the memory.
[0020] According to a third aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and when the instructions are executed by a processor, the method described in any of the above embodiments is implemented.
[0021] According to a fourth aspect of an embodiment of the present disclosure, a computer program is provided, comprising computer instructions, wherein when the computer instructions are executed by a processor, the method described in any one of the above embodiments is implemented.
[0022] Other features and advantages of the present disclosure will become apparent from the following detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0024] FIG1 is a schematic diagram of a process of a risk detection method according to an embodiment of the present disclosure;
[0025] FIG2 is a schematic diagram of a linear chain conditional random field according to an embodiment of the present disclosure;
[0026] FIG3 is a schematic diagram of a flow chart of a risk detection method according to another embodiment of the present disclosure;
[0027] FIG4 is a schematic structural diagram of a risk detection device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0028] The following will be combined with the drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments. The following description of at least one exemplary embodiment is actually only illustrative and is in no way intended to limit the present disclosure and its application or use. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present disclosure.
[0029] Unless specifically stated otherwise, the relative arrangement of components and steps, the numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present disclosure.
[0030] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.
[0031] Technologies, methods and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, such technologies, methods and equipment should be considered part of the authorization specification.
[0032] In all examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not limiting. Therefore, other examples of the exemplary embodiments may have different values.
[0033] It should be noted that like reference numerals and letters refer to like items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.
[0034] The inventors noted that in the related art, import risk detection of edible commodities is mainly carried out manually. Since the risk detection results are easily affected by human factors, it is impossible to obtain objective detection results.
[0035] Accordingly, the present disclosure provides a risk detection method that can obtain objective and accurate risk detection results.
[0036] Figure 1 is a flow chart of a risk detection method according to an embodiment of the present disclosure. In some embodiments, the following risk detection method is performed by a risk detection device.
[0037] In step 101 , ingredient content information of edible products included in customs clearance data is identified to obtain a plurality of ingredient names and corresponding content data.
[0038] In some embodiments, identifying ingredient content information of edible products included in customs clearance data includes the following steps:
[0039] 1) Processing the ingredient content information using a first machine learning model to obtain embedded information.
[0040] For example, the first machine learning model is the BERT (bidirectional encoder representations from transformers) model.
[0041] It should be noted here that bidirectional language models are better at understanding language context than unidirectional language models. The Transformer includes an encoder that reads text input and a decoder that generates predictions for the task, while the BERT model only requires an encoder mechanism.
[0042] The BERT model uses two pre-training tasks during training: Masked Language Model (LM) and Next Sentence Prediction (NSP). For example, during the BERT masking process, 15% of the words are masked, and the model is then asked to predict the masked words. Another NSP task is to predict the relationship between the previous and next sentences, predicting whether two text segments appear consecutively in the original text.
[0043] The input of the Bert model is component content information, which is converted into text embeddings through the Bert model.
[0044] 2) Processing the embedded information using a second machine learning model to obtain a feature vector.
[0045] In some embodiments, the second machine learning model is a BILSTM (Bi-directional Long Short-Term Memory) model.
[0046] It's important to note that the inputs of traditional neural networks are relatively independent, making these models suitable for scenarios where the input data isn't closely related. However, in some scenarios, such as text data, the first word input is related to the next word. For these scenarios, where prior information must be considered, RNNs (recurrent neural networks) are suitable. However, as RNNs continue to grow over time, if the time interval between t1 and t0 is large, the memory may have lost the information learned up to t0. The LSTM (long short-term memory) model was designed to address this long-term dependency problem and avoids it through deliberate design.
[0047] Sometimes, predictions may need to be determined by combining several preceding and subsequent inputs for greater accuracy. Therefore, bidirectional recurrent neural networks (BILSTMs) were proposed. BILSTMs are an extension of LSTMs. BILSTMs train two LSTMs on the input sequence (two layers side by side, with the input sequence fed intact into the first layer and then reverse-copied and passed to the second layer), enabling more comprehensive learning.
[0048] The BILSTM model obtains the embeddings generated by the Bert model and outputs a feature vector.
[0049] 3) processing the feature vector using a third machine learning model to obtain a recognition result, wherein the recognition result includes multiple component names and multiple content data;
[0050] In some embodiments, the third machine learning model is a CRF (conditional random field) model.
[0051] For example, as shown in Figure 2, assuming that the graph G contains multiple maximum clusters C, the joint probability distribution P(Y) can be decomposed into a function ψ of the random variables on these maximum clusters c (Y C ) in the form of continuous multiplication
[0052] Satisfying the formula P(Y v |X,Y w )=P(Y v |X,Y U ) is true for any node v, where w is not equal to v, and u is all nodes connected to v. In part-of-speech tagging, the main processing is the conditional random field on the linear chain, such as X = {x1, x2, ..., x n},Y={y1,y2,…,y n}, then P(yi |X,y1,y2,…,y i-1 ,y i+1 ,…,y n )=P(y i |X,y i-1 ,y i+1 ).
[0053] Here, the feature vector output by the BILSTM model is input into the CRF model. To maximize the posterior probability, we need to find an optimal path (the selection of labels corresponding to the characters). This is typically done through dynamic programming to find subpaths and ultimately output the corresponding "path" (the optimal label combination).
[0054] 4) Optimize the correspondence between multiple ingredient names and multiple content data.
[0055] The optimization process is described below through specific embodiments.
[0056] Example 1
[0057] For example, the ingredient content information includes "Gardenia Yellow 3.5%, Caramel 4%, Azurite Gum Content 70%". If "Gardenia Yellow" is not recognized, a mismatch between the ingredient name and content data will occur, as shown in Table 1.
[0058] Table 1
[0059] To avoid mismatches, punctuation marks are added after each ingredient name and each content data in the recognition results. If the first information to be processed is one of the ingredient name and content data, and the second information following the first information is the other of the ingredient name and content data, the punctuation mark after the second information is used as the grouping boundary, and the first information and the second information are matched.
[0060] It should also be noted that if each component name and each content data is followed by a punctuation mark in the recognition result, in this case, there is no need to perform the operation of adding punctuation marks, and the existing punctuation marks can be directly used for grouping and demarcation.
[0061] For example, if the first information to be processed is the ingredient name "Gardenia Yellow", and the second information after the first information is the content data "3.5%", then "Gardenia Yellow" and "3.5%" will be matched. Next, if the first information to be processed is the ingredient name "Caramel", and the second information after the first information is the content data "4%", then "Caramel" and "4%" will be matched. Next, if the first information to be processed is the ingredient name "Azurite Gum", and the second information after the first information is the content data "70%", then "Azurite Gum" and "70%" will be matched. The matching results are shown in Table 2. As can be seen from Table 2, the mismatch problem in Table 1 above has been effectively solved.
[0062] Table 2
[0063] For another example, as shown in Table 3, if the first information to be processed is the content data "4%", and the second information following the first information is the ingredient name "caramel", then "4%" and "caramel" are matched.
[0064] Table 3
[0065] It can be seen from Tables 2 and 3 above that the mismatch between ingredient names and content data can be effectively resolved through optimization processing.
[0066] Example 2
[0067] For example, if the ingredient content information includes "chicken and beef (chicken 35%, beef 65%) 15%", it is difficult to determine which ingredient name the content data "15%" after the brackets corresponds to, so hierarchical extraction errors are likely to occur.
[0068] In order to solve this problem, when there are brackets in the recognition results, and there are ingredient names inside the brackets, and the two sides outside the brackets have ingredient names and content data adjacent to the brackets, the contents inside the brackets are identified, and a correspondence between the ingredient names and content data adjacent to the brackets is established.
[0069] For example, first identify the content in the brackets, and the corresponding relationship obtained is "(chicken, 35%)", "(beef, 65%)", and then correspond the ingredient name "chicken and beef" adjacent to the brackets with the content data "15%". The results are shown in Table 4.
[0070] Table 4
[0071] It should be noted that the ingredient content information includes content such as "calcium carbonate (15%)," which does not include a hierarchical structure. To avoid processing errors, in this embodiment, the ingredient name is defined within the brackets, and the adjacent ingredient names and content data are defined on both sides of the brackets. This ensures that content with a hierarchical structure is correctly processed.
[0072] Example 3
[0073] For example, the ingredient content information includes "meat floss 10% to 20%". In this case, it is easy to cause the problem of only matching the ingredient name "meat floss" with the content data "10%", but failing to match the content data "20%" with the ingredient name "meat floss".
[0074] In order to solve this problem, when the recognition result includes a connector, and the left end of the connector has first content data and the right end of the connector has second content data, first identification information and second identification information are generated, wherein the first identification information includes information that the component name associated with the connector is greater than the first content data, and the second identification information includes information that the component name associated with the connector is less than the second content data.
[0075] For example, when the ingredient content information includes "meat floss 10%~20%", the ingredient content information includes a connector "~", and the two ends of the connector have content data "10%" and "20%" respectively. In this case, the ingredient content information is decomposed into "meat floss>10%" and "meat floss<20%" to obtain the accurate content range, as shown in Table 5.
[0076] Table 5
[0077] Example 4
[0078] Similar to the situation in Example 1, for example, if the ingredient content information includes "Gardenia Yellow 3.5%, Caramel 4%, Azure Gum Content 70%", if "Gardenia Yellow" is not recognized, there will be a mismatch between the ingredient name and content data, as shown in Table 1.
[0079] To address this issue, punctuation marks are added after each component name and each content data in the recognition results. Next, punctuation marks are used as grouping boundaries to obtain multiple groups. If a group only includes one component name, the predetermined content data is added after the component name; if a group only includes one content data, the predetermined component name is added before the content data. Then, if, in adjacent first and second groups, the first group includes the first component name and predetermined content data, and the second group includes the predetermined component name and second content data, the first and second groups are merged, and the predetermined component name and predetermined content data are deleted, so that a corresponding relationship is established between the first component name and the second content data.
[0080] For example, in the recognition results, punctuation marks are added after each ingredient name and each content data. Next, the punctuation marks are used as grouping boundaries to obtain multiple groups, as shown in the processing result 1 in Table 6.
[0081] Next, the first group contains only one ingredient named "Gardenia Yellow," so the predetermined content data is added after the ingredient name, i.e., "(Gardenia Yellow, =, 100%)." The second group adjacent to the first group contains only one content data, "3.5%." So the predetermined ingredient name is added before the content data, i.e., (null, =, 3.5%). And so on. The resulting processing result is shown in Processing Result 2 in Table 6.
[0082] Next, if the first group contains the content data "100%" and the adjacent second group contains the predetermined component name "null", the first group and the second group are merged to obtain a new group containing "(Gardenia Yellow, =, 3.5%)". The processing results are shown in the optimization results in Table 6.
[0083] Table 6
[0084] It can be seen that the above optimization processing can effectively solve the mismatch between ingredient names and content data.
[0085] In step 102, a plurality of ingredient names and corresponding content data are compared with a predetermined food safety standard library to determine a first risk value of the edible product.
[0086] It should be noted here that the obtained ingredient name and corresponding content data are compared with the food safety standard library to determine whether the content data corresponding to the ingredient name is within a predetermined range, and the first risk value of the edible product is determined based on the judgment result.
[0087] In some embodiments, first risk statement information is generated for the edible commodity, wherein the first risk statement information includes an ingredient whose content does not meet requirements of a food safety standard library.
[0088] For example, if the ingredient "black bean red" in the edible product exceeds the standard, the first risk description information includes: "(ingredient exceeded: 'black bean red')".
[0089] In step 103, a second risk value of the plurality of pieces of access information of the edible commodity included in the customs clearance data is determined based on the food access knowledge base.
[0090] In some embodiments, the access information includes whether the edible commodity has exceeded its shelf life, and at least one of the following: whether the country of origin of the edible commodity is allowed, whether the edible commodity has a certificate of origin, whether the consumer use unit of the edible commodity has an import license, whether the manufacturer of the edible commodity is validly registered, whether the manufacturer of the edible commodity has a production and processing license, and whether the manufacturer of the edible commodity has an inspection and quarantine certificate.
[0091] It should be noted that since the judgment results of the above-mentioned multiple access information are only 0 or 1, for example, for import licenses, there are only two conditions: the existence of an import license and the absence of an import license. Therefore, this access information can accurately reflect the risk of edible products. In addition, the above-mentioned multiple access information has a decisive effect on the risk of edible products. That is, as long as any one of the access information does not meet the conditions, the edible product is determined to be risky. Therefore, the above-mentioned multiple access information is directly used to judge the risk of edible products.
[0092] In some embodiments, second risk description information for the edible commodity is generated, wherein the second risk description information includes access information that does not meet requirements of a food access knowledge base.
[0093] For example, if the consumer use unit of the edible product does not have an import license, the second risk description information includes: "(the consumer use unit does not have an import license)".
[0094] In some embodiments, a first risk vector is generated, wherein the i-th element in the first risk vector corresponds to the risk value of the i-th access information in the multiple access information, 1≤i≤N, and N is the total number of the multiple access information.
[0095] It should be noted here that, by generating the first risk vector, it is possible to quickly understand which of the plurality of access information does not meet the requirements based on the first risk vector.
[0096] For example, the first risk vector is [0,1,0,0,0,0,0]. Since the second element in the first risk vector corresponds to whether the consumer use unit of the edible commodity has an import license, the first risk vector can be used to quickly determine whether the consumer use unit of the edible commodity does not have an import license.
[0097] In step 104, historical risk information of the edible commodity is obtained.
[0098] In some embodiments, the historical risk information includes at least one of the following: the historical risk of the country of origin of the edible commodity, the historical risk of the manufacturer of the edible commodity, the historical risk of the edible commodity, the packaging discreteness of the edible commodity, the historical number of order changes of the declaring unit of the edible commodity, and the degree of expiration of the edible commodity.
[0099] For example, the historical risk of the country of origin of edible commodities is calculated based on the historical data of black samples of the country of origin of edible commodities, the historical risk of the production enterprise of edible commodities is calculated based on the historical data of black samples of edible commodities, the historical risk of edible commodities is calculated based on the historical data of black samples of edible commodities, the packaging dispersion of edible commodities is calculated based on the packaging type of edible commodities, the historical number of order changes of the declaring unit of edible commodities is calculated, and the degree of distance to expiration of edible commodities is calculated based on the length of time from the expiration of edible commodities.
[0100] It's important to note that, unlike the multiple access information mentioned above, the historical risk information involved here is a non-deterministic variable, so it doesn't significantly compete for weight during supervised learning. If any single access information is included, it will overshadow the remaining risk information, and the final model will inevitably be inaccurate.
[0101] In step 105, a third risk value of the edible commodity is determined based on the historical risk information;
[0102] In some embodiments, an ensemble learning model (e.g., a random forest model) is used to process the historical risk information, and the historical risk information of the six dimensions is integrated through supervised learning to ultimately obtain a third risk value. For example, the third risk value is between 0 and 1.
[0103] It should be noted that since black samples may be determined by multiple variables, only the third risk value is output here, and the specific risk cause is not output.
[0104] In step 106, the import risk value of the edible commodity is determined based on the first risk value, the second risk value, and the third risk value.
[0105] In some embodiments, the import risk value of the edible commodity is determined based on the sum of the first risk value, the second risk value, and the third risk value.
[0106] For example, if the first risk value is 1, the second risk value is 1, and the third risk value is 0.00212802, then the import risk value of edible goods is 2.00212802.
[0107] In some embodiments, the first risk value and the third risk value are inserted into the first risk vector to obtain a second risk vector. Thus, based on the second risk vector, it is convenient to understand which dimension of the edible product information does not meet the requirements.
[0108] In some embodiments, risk cause information of the edible commodity is generated based on the first risk description information and the second risk description information, thereby conveniently understanding the specific causes of the import risk.
[0109] FIG3 is a flow chart of a risk detection method according to another embodiment of the present disclosure. As shown in FIG3 , the ingredient content judgment module identifies the ingredient content information of edible commodities included in the customs clearance data to obtain multiple ingredient names and corresponding content data. The multiple ingredient names and corresponding content data are compared with a predetermined food safety standard library to determine the first risk value of the edible commodity and the name of the ingredient exceeding the standard. The access information judgment and analysis module determines the second risk value and risk vector of the multiple access information of edible commodities included in the customs clearance data based on the food access knowledge base. Historical risk information is obtained by performing historical risk modeling on historical data. The risk information analysis module determines the third risk value of the edible commodity based on the historical risk information. The import risk value of the edible commodity is determined based on the first risk value, the second risk value, and the third risk value. The first risk value and the third risk value are inserted into the risk vector to obtain an updated risk vector. The risk reason information of the edible commodity is generated based on the first risk description information and the second risk description information.
[0110] For example, the import risk value of a food product is:
[0111] [2.00212802]
[0112] The updated risk vector is:
[0113] [1,0,1,0,0,0,0,0,0.00212802]
[0114] Since the first element of the risk vector corresponds to whether the ingredients exceed the standard, the third element corresponds to whether there is an import license, and the ninth element corresponds to whether there is a risk in the historical information, it can be seen that the import risk value of the edible product comes from the excessive ingredients, the lack of an import license, and the risk of historical information.
[0115] The corresponding risk cause information includes:
[0116] [Ingredients exceeding the standard: 'Black Bean Red', the consumer unit does not have a valid import license]
[0117] By implementing the risk detection method of the above-mentioned embodiment of the present disclosure, the import risk of edible goods can be conveniently and accurately assessed, effectively saving labor costs.
[0118] FIG4 is a schematic diagram of the structure of a risk detection device according to an embodiment of the present disclosure. As shown in FIG4 , the risk detection device includes a memory 41 and a processor 42 .
[0119] The memory 41 is used to store instructions. The processor 42 is coupled to the memory 41 . The processor 42 is configured to execute the method involved in any embodiment in FIG. 1 based on the instructions stored in the memory.
[0120] As shown in Figure 4, the risk detection device also includes a communication interface 43 for exchanging information with other devices. The risk detection device also includes a bus 44 through which the processor 42, the communication interface 43, and the memory 41 communicate with each other.
[0121] Memory 41 may include high-speed RAM memory or non-volatile memory, such as at least one disk drive. Memory 41 may also be a memory array. Memory 41 may also be divided into blocks, and the blocks may be combined into virtual volumes according to certain rules.
[0122] Furthermore, the processor 42 may be a central processing unit (CPU), or may be an application-specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the present disclosure.
[0123] The present disclosure also relates to a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and when the instructions are executed by a processor, the method involved in any embodiment of FIG. 1 is implemented.
[0124] By implementing the above embodiments of the present disclosure, the following beneficial effects can be achieved:
[0125] 1. This paper proposes a model method based on supervised learning, which introduces machine learning into risk diagnosis, quantitatively assesses the risk level of edible products, facilitates real-time judgment, and saves labor costs.
[0126] 2. The present disclosure integrates multiple factors into the scenario of joint judgment in integrated learning, which increases the possibility of actual deployment scenarios.
[0127] 3. The present disclosure adopts a method combining framework and logic optimization in the process of identifying component content information, so that the language model can output accurate results while maintaining its interpretability.
[0128] 4. This disclosure not only provides risk scores, but also provides risk cause information, making the output results easy to understand and able to quickly locate the root cause of the problem.
[0129] In some embodiments, the functional units described above may be implemented as general-purpose processors, programmable logic controllers (PLC), digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or any appropriate combination thereof, for performing the functions described in the present disclosure.
[0130] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.
[0131] The description of the present disclosure is provided for purposes of illustration and description and is not intended to be exhaustive or to limit the disclosure to the disclosed form. Many modifications and variations will be apparent to those skilled in the art. The embodiments are selected and described in order to better illustrate the principles and practical applications of the present disclosure and to enable those skilled in the art to understand the present disclosure and design various embodiments with various modifications suitable for specific applications.
Claims
1. A risk detection method, comprising: Identifying the ingredient content information of the edible goods included in the customs clearance data to obtain multiple ingredient names and corresponding content data; Comparing the multiple ingredient names and corresponding content data with a predetermined food safety standard library to determine the first risk value of the edible goods; Determining the second risk value of multiple admission information of the edible goods included in the customs clearance data according to the food admission knowledge base; Obtaining the historical risk information of the edible goods; Determining the third risk value of the edible goods according to the historical risk information; Determining the import risk value of the edible goods according to the first risk value, the second risk value and the third risk value.
2. The risk detection method according to claim 1, wherein, The determining the import risk value of the edible goods includes: Determining the import risk value of the edible goods according to the sum of the first risk value, the second risk value and the third risk value.
3. The risk detection method according to claim 1, wherein The identifying the ingredient content information of the edible goods included in the customs clearance data includes: Processing the ingredient content information by using a first machine learning model to obtain embedding information; Processing the embedding information by using a second machine learning model to obtain a feature vector; Processing the feature vector by using a third machine learning model to obtain an identification result, wherein the identification result includes the multiple ingredient names and multiple content data; Optimizing the correspondence relationship between the multiple ingredient names and multiple content data.
4. The risk detection method according to claim 3, wherein, The optimizing the correspondence relationship between the multiple ingredient names and multiple content data includes: In the identification result, adding punctuation marks after each ingredient name and each content data; If the current first information to be processed is one of the ingredient name and the content data, and the second information after the first information is the other of the ingredient name and the content data, then using the punctuation mark after the second information as a grouping boundary and corresponding the first information and the second information.
5. The risk detection method according to claim 3, wherein The optimizing the correspondence relationship between the multiple ingredient names and multiple content data includes: In the case that there are brackets in the identification result, and there are ingredient names inside the brackets, and there are an ingredient name and content data adjacent to both sides outside the brackets respectively, identifying the content inside the brackets and establishing a correspondence relationship with the ingredient name and content data adjacent to the brackets.
6. The risk detection method according to claim 3, wherein, The optimizing the correspondence relationship between the multiple ingredient names and multiple content data includes: In the case that the identification result includes a connector, and the left end of the connector has a first content data and the right end of the connector has a second content data, generating a first identification information and a second identification information, wherein the first identification information includes the information that the ingredient name associated with the connector is greater than the first content data, and the second identification information includes the information that the ingredient name associated with the connector is less than the second content data.
7. The risk detection method according to claim 3, wherein, The optimizing the correspondence relationship between the multiple ingredient names and multiple content data includes: In the identification result, adding punctuation marks after each ingredient name and each content data; Use the punctuation mark as a grouping boundary to obtain multiple groups. Among them, if a group only includes one ingredient name, add predetermined content data after the ingredient name. If a group only includes one content data, add a predetermined ingredient name before the content data; If in adjacent first and second groups, the first group includes a first ingredient name and the predetermined content data, and the second group includes the predetermined ingredient name and a second content data, then merge the first group and the second group, and delete the predetermined ingredient name and the predetermined content data, so that the first ingredient name and the second content data establish a corresponding relationship.
8. The risk detection method according to claim 3, wherein, The first machine learning model is a BERT model; The second machine learning model is a BILSTM model; The third machine learning model is a CRF model.
9. The risk detection method according to claim 1, wherein, The access information includes whether the edible commodity has exceeded the shelf life, and at least one of the following: whether the country of origin of the edible commodity is admissible, whether the edible commodity has a certificate of origin, whether the consumer and user unit of the edible commodity has an import license, whether the production enterprise of the edible commodity is effectively registered, whether the production enterprise of the edible commodity has a production and processing license, and whether the production enterprise of the edible commodity has an inspection and quarantine certificate.
10. The risk detection method according to claim 1, wherein, The historical risk information includes at least one of the following: the historical risk of the country of origin of the edible commodity, the historical risk of the production enterprise of the edible commodity, the historical risk of the commodity of the edible commodity, the packaging dispersion of the edible commodity, the historical number of amendment requests of the declaration unit of the edible commodity, and the degree of proximity to expiration of the edible commodity.
11. The risk detection method according to claim 10, wherein, Determining the third risk value of the edible commodity according to the historical risk information includes: Using an ensemble learning model to process the historical risk information to obtain the third risk value.
12. The risk detection method according to any one of claims 1-11 further includes: Generate a first risk vector, where the i-th element in the first risk vector corresponds to the risk value of the i-th access information in the multiple access information, 1≤i≤N, and N is the total number of the multiple access information; Insert the first risk value and the third risk value into the first risk vector to obtain a second risk vector.
13. The risk detection method according to claim 12 further includes: Generate first risk description information of the edible commodity, where the first risk description information includes ingredients whose content does not meet the requirements of the food safety standard library; Generate second risk description information of the edible commodity, where the second risk description information includes access information that does not meet the requirements of the food access knowledge base; Generate risk cause information of the edible commodity according to the first risk description information and the second risk description information.
14. A risk detection device, comprising: A memory; A processor, coupled to a memory, the processor being configured to execute a risk detection method as described in any one of claims 1-13 based on instructions stored in the memory.
15. A computer-readable storage medium, wherein, A computer-readable storage medium stores computer instructions, which when executed by a processor implement a risk detection method as described in any one of claims 1-13.
16. A computer program, comprising computer instructions, wherein the computer instructions, when executed by a processor, implement a risk detection method as described in any one of claims 1-13.
Citation Information
Patent Citations
Cross-border e-commerce commodity quality risk identification method based on rules
CN107886240A
Automatic grading and early warning system and method for safety risk of imported food
CN113762764A
Food risk comprehensive evaluation method based on improved matter element extension model
CN113887978A
Food safety risk prediction method and system
CN115964504A
Risk detection method and device and storage medium
CN117875984A