Label determination method and device for machine account data, equipment and medium

By classifying and labeling data of security supervision objects, the problem of poor accuracy in labeling labeling in traditional technology is solved, and higher label information reliability and business needs adaptability are achieved.

CN120179827AActive Publication Date: 2025-06-20北京市应急指挥保障中心
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510670623.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-20
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

Traditional regulatory systems have poor accuracy in determining ledger data labels for safety supervision objects, which is difficult to meet the multi-dimensional and dynamic risk warning and emergency command needs.

Method used

By obtaining the ledger data of the security supervision objects to be processed, using preset classification rules for classification processing, and combining the preset marking strategy and label prediction model, the label information of the ledger data is determined.

Benefits of technology

It improves the reliability and accuracy of ledger data label information, and can mark the data in combination with data types to meet dynamic business needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179827A_ABST
    Figure CN120179827A_ABST
Patent Text Reader

Abstract

The invention discloses a label determination method and device for machine account data, equipment and a medium, and relates to the technical field of computers. The method comprises the following steps: acquiring standing book data of a to-be-processed security supervision object; utilizing a preset classification rule to perform classification processing on the machine account data so as to obtain a first type of machine account data and a second type of machine account data; marking the first type of machine account data by using a preset marking strategy to obtain label information of the first type of machine account data; inputting the second type of machine account data into a preset label prediction model to obtain label information of the second type of machine account data; and determining the label information of the machine account data based on the label information of the first type of machine account data and the label information of the second type of machine account data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, specifically to technical fields such as big data technology and emergency safety management technology, and particularly relates to a method, device, equipment, and medium for determining labels of ledger data. Background Art

[0002] In the field of emergency safety management, the basic information collection mechanism of traditional supervision systems has significant technical limitations. Specifically, after the initial collection of enterprise entity information and site basic data and the formation of ledger data of safety supervision objects, as the safety production supervision business develops in a multi-dimensional and dynamic direction, the original data structure gradually becomes difficult to meet the functional evolution requirements of business systems such as risk warning and emergency command, and it is necessary to expand the attributes of the collected data.

[0003] Currently, the relevant technologies mainly include solutions for heterogeneous data fusion governance and manual annotation extension mechanisms. The solution for heterogeneous data fusion governance is to collect more data from other business systems associated with the enterprise and, through big data governance, form complex basic databases and label databases to solve the problem. The solution for the manual annotation extension mechanism is to form extended attributes by simply manually tagging the basic information. Summary of the Invention

[0004] This application provides a method, device, equipment, and medium for determining labels of ledger data, which can solve the problem of poor accuracy in determining labels of ledger data. The technical solutions are as follows: In the first aspect, a method for determining labels of ledger data is provided. The method includes: Obtain the ledger data of the safety supervision object to be processed; Use a preset classification rule to classify the ledger data to obtain first-category ledger data and second-category ledger data; Use a preset tagging strategy to tag the first-category ledger data to obtain the label information of the first-category ledger data; Input the second-category ledger data into a preset label prediction model to obtain the label information of the second-category ledger data; Based on the label information of the first-category ledger data and the label information of the second-category ledger data, determine the label information of the ledger data.

[0005] In a possible implementation manner, the step of using a preset classification rule to classify the ledger data to obtain first-category ledger data and second-category ledger data includes: In response to the ledger data including not only the entity type to which the ledger belongs, the organization to which the ledger belongs, the region to which the ledger belongs, and the ledger text description, determining the ledger data as the first category of ledger data; In response to the ledger data only including the entity type to which the ledger belongs, the organization to which the ledger belongs, the region to which the ledger belongs, and the ledger text description, determining the ledger data as the second category of ledger data.

[0006] In a possible implementation manner, the using a preset tagging strategy to perform a tagging process on the first category of ledger data to obtain the tag information of the first category of ledger data includes: Using a preset attribute tagging rule to perform a tagging process on the first category of ledger data to obtain the attribute tag corresponding to the first category of ledger data; Using a preset business tagging rule to perform a tagging process on the first category of ledger data to obtain the business tag corresponding to the first category of ledger data; Based on the attribute tag corresponding to the first category of ledger data, and / or, the business tag corresponding to the first category of ledger data, obtaining the tag information of the first category of ledger data.

[0007] In a possible implementation manner, the second category of ledger data includes the entity type to which the ledger belongs, the organization to which the ledger belongs, the region to which the ledger belongs, and the ledger text description, the preset tag prediction model includes a fully connected network, and the inputting the second category of ledger data into the preset tag prediction model to obtain the tag information of the second category of ledger data includes: Inputting the entity type to which the ledger belongs, the organization to which the ledger belongs, the region to which the ledger belongs, and the ledger text description into the preset tag prediction model; Respectively performing an encoding process on the entity type to which the ledger belongs, the organization to which the ledger belongs, the region to which the ledger belongs, and the ledger text description to obtain a type encoding feature, an organization encoding feature, a region encoding feature, and a text encoding feature; Performing a splicing process on the type encoding feature, the organization encoding feature, the region encoding feature, and the text encoding feature to obtain a fusion feature; Inputting the fusion feature into the fully connected network to obtain a feature vector corresponding to the fusion feature; Based on the feature vector corresponding to the fusion feature, obtaining the tag information of the second category of ledger data from a preset tag feature vector database.

[0008] In a possible implementation, the preset label prediction model further includes a graph attention encoding network and a text embedding encoding network, which respectively perform encoding processing on the entity type to which the ledger belongs, the organization to which the ledger belongs, the region to which the ledger belongs, and the ledger text description to obtain a type encoding feature, an organization encoding feature, a region encoding feature, and a text encoding feature, including: Respectively use the graph attention encoding network to perform encoding processing on the entity type to which the ledger belongs, the organization to which the ledger belongs, and the region to which the ledger belongs to obtain the type encoding feature, the organization encoding feature, and the region encoding feature; Use the text embedding encoding network to perform encoding processing on the ledger text description to obtain a text encoding feature.

[0009] In a possible implementation, before inputting the second category of ledger data into the preset label prediction model to obtain the label information of the second category of ledger data, it includes: Obtain sample ledger data and the label information corresponding to the sample ledger data; the sample ledger data includes the entity type to which the sample ledger belongs, the organization to which the sample ledger belongs, the region to which the sample ledger belongs, and the sample ledger text description; Input the entity type to which the sample ledger belongs, the organization to which the sample ledger belongs, the region to which the sample ledger belongs, the sample ledger text description, and the label information corresponding to the sample ledger data into the label prediction model to be trained; Respectively perform encoding processing on the entity type to which the sample ledger belongs, the organization to which the sample ledger belongs, the region to which the sample ledger belongs, and the sample ledger text description to obtain a type encoding feature, an organization encoding feature, a region encoding feature, and a text encoding feature; Perform splicing processing on the type encoding feature, the organization encoding feature, the region encoding feature, and the text encoding feature to obtain a fusion feature; Input the fusion feature into the fully connected network of the label prediction model to be trained to obtain a feature vector corresponding to the fusion feature; Perform encoding processing on the label information corresponding to the sample ledger data to obtain a label feature vector; Based on the label feature vector, the feature vector corresponding to the fusion feature, the sample ledger data, and the label information, update and train the label prediction model to be trained to obtain a trained label prediction model.

[0010] In a second aspect, a device for determining labels of ledger data is provided, and the device includes: An acquisition unit for acquiring ledger data of a safety supervision object to be processed; A classification unit for classifying the ledger data by using a preset classification rule to obtain first-category ledger data and second-category ledger data; A marking unit for marking the first-category ledger data by using a preset marking strategy to obtain label information of the first-category ledger data; A prediction unit for inputting the second-category ledger data into a preset label prediction model to obtain label information of the second-category ledger data; A determination unit for determining label information of the ledger data based on the label information of the first-category ledger data and the label information of the second-category ledger data.

[0011] In a third aspect, a computer-readable storage medium is provided, in which at least one instruction is stored, and the at least one instruction is loaded and executed by a processor to implement the method in the above aspect and any possible implementation manner.

[0012] In a fourth aspect, an electronic device is provided, including: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method in the above aspect and any possible implementation manner.

[0013] In a fifth aspect, a computer program product is provided, including a computer program which, when executed by a processor, implements the method in the above aspect and any possible implementation manner.

[0014] The beneficial effects of the technical solution provided in this application at least include: As can be seen from the above technical solution, in the embodiment of this application, the ledger data of the safety supervision object to be processed can be obtained, and then the ledger data can be classified by using a preset classification rule to obtain first-category ledger data and second-category ledger data. The first-category ledger data is marked by using a preset marking strategy to obtain label information of the first-category ledger data. The second-category ledger data is input into a preset label prediction model to obtain label information of the second-category ledger data. The label information of the ledger data is determined based on the label information of the first-category ledger data and the label information of the second-category ledger data. Since the label information of the ledger data can be determined by using a preset marking strategy and a preset label prediction model respectively, and the marking process can be carried out separately according to the type of the ledger data, the reliability and accuracy of the label information of the ledger data are improved.

[0015] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0017] Figure 1 is a schematic flowchart of a method for determining labels of ledger data provided by an embodiment of the present application; Figure 2 is a schematic diagram of a preset label prediction model in the method for determining labels of ledger data provided by an embodiment of the present application; Figure 3 is a schematic diagram of a label prediction model to be trained in the method for determining labels of ledger data provided by an embodiment of the present application; Figure 4 is a structural block diagram of a device for determining labels of ledger data provided by an embodiment of the present application; Figure 5 is a block diagram of an electronic device for implementing the method for determining labels of ledger data in the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] The following will describe exemplary embodiments of the present application with reference to the accompanying drawings, including various details of the embodiments of the present application to facilitate understanding. It should be considered that they are merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for clarity and conciseness, the description of well-known functions and structures is omitted below.

[0019] Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.

[0020] It should be noted that the terminal devices involved in the embodiments of the present application may include, but are not limited to, intelligent devices such as mobile phones, personal digital assistants (PDAs), wireless handheld devices, and tablet computers; the display devices may include, but are not limited to, devices with display functions such as personal computers and televisions.

[0021] In addition, the term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.

[0022] Currently, for the heterogeneous data fusion governance method, the processing cycle of label data is relatively long, usually taking several weeks to several months to supplement quantitative parameters, unable to respond in real time to the new data feature requirements in emergency incident handling, nor can it provide the new features required by the new business system in a timely manner. The simple manual labeling mode also has three disadvantages: easy to make mistakes, large workload, and inconvenient management.

[0023] Therefore, there is an urgent need for a method for determining the labels of ledger data to effectively perform labeling processing on ledger data, thereby ensuring the reliability and accuracy of the label information of ledger data.

[0024] Please refer to Figure 1 , which shows a schematic flowchart of a method for determining the labels of ledger data provided by an embodiment of the present application. The method for determining the labels of ledger data may specifically include: Step 101, obtain the ledger data of the safety supervision object to be processed.

[0025] Step 102, use the preset classification rules to classify the ledger data to obtain the first-category ledger data and the second-category ledger data.

[0026] Step 103, use the preset labeling strategy to perform labeling processing on the first-category ledger data to obtain the label information of the first-category ledger data.

[0027] Step 104, input the second-category ledger data into the preset label prediction model to obtain the label information of the second-category ledger data.

[0028] Step 105, based on the label information of the first-category ledger data and the label information of the second-category ledger data, determine the label information of the ledger data.

[0029] It should be noted that the ledger data of the safety supervision objects can be the basic information of the safety supervision objects collected in real time or regularly from various emergency management committees, various industry units, big data centers, and basic platforms. Based on the ledger data of the safety supervision objects, a ledger database of the safety supervision objects can be constructed.

[0030] It should be noted that the ledger data of the safety supervision objects can serve emergency operations, such as the emergency safety hazard investigation business system. The emergency safety hazard investigation business system can use the ledger data and, at the same time, correct the ledger data during the investigation process and directly update the ledger database of the safety supervision objects by backflow.

[0031] It should be noted that the label information of the second category of ledger data can include multiple labels corresponding to the ledger data.

[0032] It should be noted that part or all of the execution subjects of steps 101 to 105 can be an application located on a local terminal, or can also be a plug-in or software development kit (SDK) and other functional units set in the application located on the local terminal, or can also be a processing engine located in a network-side server, or can also be a distributed system located on the network side. For example, the ledger data platform processing engine or distributed system of the safety supervision objects on the network side, etc. This embodiment does not make special limitations on this.

[0033] It can be understood that the application can be a native app installed on the local terminal, or can also be a web app of a browser on the local terminal. This embodiment does not make limitations on this.

[0034] In this way, by obtaining the ledger data of the safety supervision objects to be processed, the ledger data can be classified by using a preset classification rule to obtain the first category of ledger data and the second category of ledger data. The first category of ledger data is labeled by using a preset labeling strategy to obtain the label information of the first category of ledger data. The second category of ledger data is input into a preset label prediction model to obtain the label information of the second category of ledger data. Based on the label information of the first category of ledger data and the label information of the second category of ledger data, the label information of the ledger data is determined. Since the label information of the ledger data can be determined by using a preset labeling strategy and a preset label prediction model respectively, and the labeling process can be carried out separately according to the type of the ledger data, the reliability and accuracy of the label information of the ledger data are improved.

[0035] Optionally, in a possible implementation of this embodiment, in step 102, specifically, in response to the ledger data including not only the entity type to which the ledger belongs, the organization to which the ledger belongs, the region to which the ledger belongs, and the ledger text description, the ledger data is determined as the first category of ledger data; in response to the ledger data only including the entity type to which the ledger belongs, the organization to which the ledger belongs, the region to which the ledger belongs, and the ledger text description, the ledger data is determined as the second category of ledger data.

[0036] In this implementation, the preset classification rules may include determining whether the ledger data only includes the entity type to which the ledger belongs, the organization to which the ledger belongs, the region to which the ledger belongs, and the ledger text description.

[0037] In this implementation, the first category of ledger data may be data with relatively comprehensive attributes of the safety supervision object. The second category of ledger data may be data with partial attributes of the safety supervision object.

[0038] In this implementation, the ledger data of the safety supervision object may include, but is not limited to, the entity type to which the ledger belongs, the organization to which the ledger belongs, the region to which the ledger belongs, the ledger text description, the unit name, the address information, the person in charge information, the area information, the number of personnel, the longitude and latitude, etc.

[0039] Here, the entity type to which the ledger belongs may include the unified credit code, the affiliated industry management department, the industry category, the industry classification, etc. The organization to which the ledger belongs may include the affiliated community information, the affiliated supervision department, etc. The region to which the ledger belongs may include the address information, the district and street information. The ledger text description may include description information such as the main business, remarks, and other businesses.

[0040] Optionally, in a possible implementation of this embodiment, in step 103, specifically, the preset attribute label rules may be used to label the first category of ledger data to obtain the attribute labels corresponding to the first category of ledger data, and the preset business label rules may be used to label the first category of ledger data to obtain the business labels corresponding to the first category of ledger data. Based on the attribute labels corresponding to the first category of ledger data, and / or, the business labels corresponding to the first category of ledger data, the label information of the first category of ledger data is obtained.

[0041] In a specific implementation process of this implementation, based on the first category of ledger data, the attribute labels corresponding to the first category of ledger data are obtained from the preset relationship table between the attribute labels and the ledger data; In another specific implementation process of this implementation, based on the first category of ledger data, the business labels corresponding to the first category of ledger data are obtained from the preset relationship table between the business labels and the ledger data.

[0042] In this implementation manner, the relationship table between the preset attribute tags and the ledger data can be a correspondence table between the predefined attribute tags and the ledger data. The preset attribute tags can be tags determined based on the basic attributes of the ledger data. For example, if the ledger data includes hazardous chemicals and the location address is within a key area, the attribute tag configured for this ledger data is a major hazard source tag.

[0043] In this implementation manner, the relationship table between the preset business tags and the ledger data is a correspondence table between the predefined business tags and the ledger data. The preset business tags can be tags determined based on the safety supervision situation of the ledger data. The ledger data can also include more data such as self-inspection results, hidden danger results, and the production and operation situation of the location. For example, if the production and operation situation of a location is that hidden dangers have been found many times in a year, the corresponding business tag is a key observation object tag.

[0044] In another specific implementation process of this implementation manner, the attribute tags corresponding to the first category of ledger data can be used as the label information of the first category of ledger data.

[0045] In yet another specific implementation process of this implementation manner, the business tags corresponding to the first category of ledger data can be used as the label information of the first category of ledger data.

[0046] In yet another specific implementation process of this implementation manner, conflict filtering processing is performed on the business tags and attribute tags corresponding to the first category of ledger data, and based on the result of the conflict filtering processing, the label information of the first category of ledger data is obtained.

[0047] Optionally, in a possible implementation manner of this embodiment, the second category of ledger data may include the entity type to which the ledger belongs, the organization to which the ledger belongs, the region to which the ledger belongs, and the ledger text description. The preset label prediction model may include a fully connected network. In step 104, first, the entity type to which the ledger belongs, the organization to which the ledger belongs, the region to which the ledger belongs, and the ledger text description are input into the preset label prediction model. Second, encoding processing is respectively performed on the entity type to which the ledger belongs, the organization to which the ledger belongs, the region to which the ledger belongs, and the ledger text description to obtain a type encoding feature, an organization encoding feature, a region encoding feature, and a text encoding feature. Third, the type encoding feature, the organization encoding feature, the region encoding feature, and the text encoding feature are concatenated to obtain a fusion feature. Fourth, the fusion feature is input into the fully connected network to obtain a feature vector corresponding to the fusion feature. Fifth, based on the feature vector corresponding to the fusion feature, the label information of the second category of ledger data is obtained from the preset label feature vector database.

[0048] In this implementation manner, the preset label prediction model further includes a graph attention encoding network and a text embedding encoding network.

[0049] In a specific implementation process of this implementation manner, the graph attention encoding network can be respectively used to encode the entity type to which the ledger belongs, the organization to which the ledger belongs, and the region to which the ledger belongs, so as to obtain the type encoding feature, the organization encoding feature, and the region encoding feature. Furthermore, the text embedding encoding network can be used to encode the ledger text description to obtain the text encoding feature.

[0050] One case of this specific implementation process is to use the graph attention encoding network to encode the entity type to which the ledger belongs to obtain the type encoding feature.

[0051] One case of this specific implementation process is to use the graph attention encoding network to encode the organization to which the ledger belongs to obtain the organization encoding feature.

[0052] One case of this specific implementation process is to use the graph attention encoding network to encode the region to which the ledger belongs to obtain the region encoding feature.

[0053] One case of this specific implementation process is to use the text embedding encoding network to perform stop word removal, keyword extraction, and text encoding processing on the ledger text description to obtain the text encoding feature.

[0054] In a specific implementation process of this implementation manner, the number of fully connected networks can be 3. First, the fusion feature can be input into the first fully connected network, and the first intermediate feature is output. The first intermediate feature is input into the second fully connected network, and the second intermediate feature is output. The second intermediate feature is input into the third fully connected network, and the feature vector corresponding to the fusion feature is output.

[0055] In another specific implementation process of this implementation manner, based on the feature vector corresponding to the fusion feature, at least one label vector similar to the feature vector corresponding to the fusion feature can be found from the preset label feature vector database by using the nearest neighbor algorithm. Based on the at least one label vector, the label information of the second category of ledger data can be obtained.

[0056] In this implementation manner, the preset label feature vector database can include a faiss vector database and other vector databases. This preset label feature vector database caches the label vectors during the training of the label prediction model. The label vector is obtained by encoding the label information corresponding to the sample ledger data.

[0057] Exemplarily,Figure 2 It is a schematic diagram of a preset label prediction model in the method for determining labels of ledger data provided by an embodiment of the present application, as Figure 2 shown.

[0058] First, use the graph attention encoding network to perform graph attention encoding (graph attention encoding) on the entity type to which the ledger belongs, the organization to which the ledger belongs, and the region to which the ledger belongs, respectively, to obtain the type encoding feature (sensetype embeding), the organization encoding feature (org embedding), and the region encoding feature (region embedding), which are denoted as SE, OE, and RE respectively, and their feature dimensions are all the size of the embedding dimension (embed_dim). embed_dim can be set according to actual needs.

[0059] Second, perform text embedding encoding on the ledger text description, and respectively go through the processes of removing stop words, keyword extraction, and text encoding to obtain the text description encoding description embedding, denoted as DE, and the feature dimension size is embed_dim.

[0060] Third, splice the type encoding feature, the organization encoding feature, the region encoding feature, and the text encoding feature in the feature dimension to obtain the fusion feature [SE, OE, RE, DE], and the dimension is 4×embed_dim;

[0061] Third, the first fully connected network consists of a Linear layer and an activation layer ReLu. The fusion feature is input into the first fully connected network, and the first intermediate feature is output. The input dimension is 4×embed_dim, and the output dimension is 2×embed_dim.

[0062] Third, the first intermediate feature is input into the second fully connected network, and the second intermediate feature is output. The input dimension is 2×embed_dim, and the output dimension is embed_dim.

[0063] Third, the second intermediate feature is input into the third fully connected network to obtain the feature vector corresponding to the fusion feature, that is, the ledger feature vector. The input dimension is embed_dim, and the output dimension is embed_dim.

[0064] Third, based on the feature vector corresponding to the fusion feature, use the nearest neighbor algorithm to find multiple label vectors that are most similar to the feature vector corresponding to the fusion feature from the preset label feature vector database as recommended labels, and obtain the label information of the second category of ledger data based on the recommended labels.

[0065] Optionally, in a possible implementation manner of this embodiment, before step 104, first, sample ledger data and the label information corresponding to the sample ledger data can be obtained. The sample ledger data includes the entity type to which the sample ledger belongs, the organization to which the sample ledger belongs, the region to which the sample ledger belongs, and the text description of the sample ledger. Secondly, the entity type to which the sample ledger belongs, the organization to which the sample ledger belongs, the region to which the sample ledger belongs, the text description of the sample ledger, and the label information corresponding to the sample ledger data can be input into the label prediction model to be trained. Thirdly, the entity type to which the sample ledger belongs, the organization to which the sample ledger belongs, the region to which the sample ledger belongs, and the text description of the sample ledger can be respectively encoded to obtain type encoding features, organization encoding features, region encoding features, and text encoding features. Thirdly, the type encoding features, organization encoding features, region encoding features, and text encoding features can be concatenated to obtain a fusion feature. Thirdly, the fusion feature can be input into the fully connected network of the label prediction model to be trained to obtain a feature vector corresponding to the fusion feature. Thirdly, the label information corresponding to the sample ledger data is encoded to obtain a label feature vector. Thirdly, based on the label feature vector, the feature vector corresponding to the fusion feature, the sample ledger data, and the label information, the label prediction model to be trained is updated and trained to obtain a trained label prediction model.

[0066] In this implementation manner, the label prediction model to be trained can include a ledger encoding tower, a label encoding tower, and a normalization (softmax) layer. The ledger encoding tower can include a graph attention encoding network, a text embedding encoding network, a first fully connected network, a second fully connected network, and a third fully connected network. The label encoding tower can include a graph attention encoding network.

[0067] In a specific implementation process of this implementation manner, the label information corresponding to the sample ledger data is input into the label encoding tower, and a label feature vector is output, and the output feature dimension is .

[0068] In this implementation manner, the loss function of the softmax layer of the label prediction model to be trained can be a cross-entropy loss function: In this specific implementation process, the label information corresponding to the sample ledger data is input into the label encoding tower, and a label feature vector is output, and the output feature dimension is .

[0069] Exemplarily, Figure 3 is a schematic diagram of the label prediction model to be trained in the method for determining the label of the ledger data provided in an embodiment of the present application, as Figure 3 shown.

[0070] Here, the label prediction model to be trained may include a ledger coding tower, a label coding tower, and a normalization (softmax) layer. The ledger coding tower may include a graph attention coding network, a text embedding coding network, a first fully connected network, a second fully connected network, and a third fully connected network. First, the graph attention coding network is used to perform graph attention coding on the entity type to which the sample ledger belongs, the organization to which the sample ledger belongs, and the region to which the sample ledger belongs, respectively, to obtain a type coding feature, an organization coding feature, and a region coding feature. Secondly, the text embedding coding network is used to perform text embedding coding on the text description of the sample ledger, and through the processes of removing stop words, keyword extraction, and text coding respectively, a text description coding is obtained. Thirdly, the type coding feature, the organization coding feature, the region coding feature, and the text coding feature are concatenated in the feature dimension to obtain a fusion feature. Thirdly, the first fully connected network consists of a Linear layer and an activation layer ReLu. The fusion feature is input into the first fully connected network, and a first intermediate feature is output. Thirdly, the first intermediate feature is input into the second fully connected network, and a second intermediate feature is output. Thirdly, the second intermediate feature is input into the third fully connected network to obtain a feature vector corresponding to the fusion feature, that is, a ledger feature vector. Thirdly, the label coding tower is a graph attention coding network. The label information corresponding to the sample ledger data is input into the label coding tower, and a label feature vector is output. Thirdly, the feature vector corresponding to the fusion feature and the label feature vector are multiplied, and the result of the multiplication is input into the softmax layer to obtain a normalized probability, so as to perform update training on the label prediction model to be trained, and a trained label prediction model is obtained.

[0071] In addition, the label feature vector obtained after encoding the label information corresponding to the sample ledger data can be stored in a preset label feature vector database.

[0072] The graph attention coding network in the label prediction model to be trained includes an attention mechanism module. The information input into the graph attention coding network, for example, the label information input into the label coding tower, can be used as elements. The elements and adjacent elements are respectively embedded to obtain , , indicating the element embedding information, indicating the adjacent element embedding information, indicating the belonging relationship, indicating the number of points with an adjacent relationship. The and are input into the attention mechanism module to calculate the attention score of the adjacent information and the adjacent relationship of this point . The is transposed and multiplied by to obtain the final coding : 。

[0073] Optionally, the attention mechanism module may include two fully connected layers and a softmax layer, and input the specified element embedding information and the adjacent element embedding information , copy itself to expand the dimension of the key element embedding information to , representing the intermediate feature information, concatenate the node and adjacent node embedding information on the dimension where is located , through two fully connected layers, each fully connected layer performs activation. Through the first fully connected layer, the feature dimension changes from and the output is . The input of the second fully connected layer is , and the output is also . The function of the softmax layer calculates the final attention score to obtain the output of the attention mechanism module

[0074] It can be understood that for some ledger data with less information, a label prediction model can be trained through historical label information to recommend labels for these ledger data to achieve labeling of ledger data

[0075] Here, when defining the label prediction model represents the label set represents the number of labels. At the same time, there is a tree relationship between the labels , represents an a label set represents the connection relationship set between the a label and the label, the connection is 1, otherwise it is 0. Each ledger has attributes, that is, ledger data, which are entity classification (assuming the fixed type number is ), the organization to which it belongs (assuming the fixed number is ), the level of the geographical area to which it belongs (administrative division) (assuming the fixed number is ), and the ledger text description, expressed as: , where the geographical area level, the organization to which it belongs, and the label are similar and have a tree relationship, which are respectively defined as , represents the geographical area level set represents the connection relationship set between the geographical area levels represents the set of organizations to which it belongs represents the connection relationship set between the organizations. The prediction problem is expressed as: According to the historical prediction data tuple , represents an element represents a platform user Represents the ledger Represents the label, for the next given ledger Before prediction when marking Previous label, the mathematical definition of the label prediction model can be: . Among them, Represents the previous label results obtained by prediction Represents the mapping relationship

[0076] Here, when training the label prediction model, by designing the encoding of the model, different feature conditions can be mapped to the same space, training the discrimination degree of the distribution space features of the encoded vectors, and finally when making predictions, encoding the existing attributes, and finding similar distribution feature vectors in the encoded space as the final prediction results. When using the trained label prediction model to predict labels, input the structured features such as ledger type, affiliated organization, affiliated geographical level and unstructured feature ledger text description in the ledger data into the trained label model, perform label recommendation, and obtain the predicted recommended labels topN, and automatic marking processing can be performed. Moreover, multiple recommended labels can also be displayed to the platform users for the users to perform marking processing

[0077] In this way, based on the sample data, the label prediction model to be trained including the ledger encoding tower and the label encoding tower can be updated and trained to obtain the trained label prediction model, which can improve the performance and prediction reliability of the label prediction model

[0078] It should be noted that the specific implementation process provided in this implementation manner can be combined with the multiple specific implementation processes provided in the foregoing implementation manner to implement the method for determining the label of the ledger data in this embodiment. For a detailed description, reference can be made to the relevant content in the foregoing implementation manner, which will not be elaborated here

[0079] Figure 4 Shows the structural block diagram of the device for determining the label of the ledger data provided by an embodiment of the present application, as Figure 4As shown in the figure. The label determination device 400 for the ledger data of this embodiment may include an acquisition unit 401, a classification unit 402, a labeling unit 403, a prediction unit 404, and a determination unit 405. Among them, the acquisition unit 401 is used to acquire the ledger data of the safety supervision object to be processed; the classification unit 402 is used to classify the ledger data by using a preset classification rule to obtain the first-category ledger data and the second-category ledger data; the labeling unit 403 is used to label the first-category ledger data by using a preset labeling strategy to obtain the label information of the first-category ledger data; the prediction unit 404 is used to input the second-category ledger data into a preset label prediction model to obtain the label information of the second-category ledger data; the determination unit 405 is used to determine the label information of the ledger data based on the label information of the first-category ledger data and the label information of the second-category ledger data.

[0080] It should be noted that part or all of the label determination device for the ledger data of this embodiment may be an application located on the local terminal, or may also be a functional unit such as a plug-in or software development kit (SDK) set in the application located on the local terminal, or may also be a processing engine located in the network-side server, or may also be a distributed system located on the network side. For example, the processing engine or distributed system in the ledger data platform of the safety supervision object on the network side, etc. This embodiment does not make special limitations on this.

[0081] It can be understood that the application may be a native program (nativeApp) installed on the local terminal, or may also be a web program (webApp) of the browser on the local terminal. This embodiment does not make limitations on this.

[0082] Optionally, in a possible implementation manner of this embodiment, the classification unit 402 is used to determine that the ledger data is the first-category ledger data in response to the ledger data including not only the ledger-owned entity type, ledger-owned organization, ledger-owned region, and ledger text description; and determine that the ledger data is the second-category ledger data in response to the ledger data only including the ledger-owned entity type, ledger-owned organization, ledger-owned region, and ledger text description.

[0083] Optionally, in a possible implementation of this embodiment, the labeling unit 403 is configured to perform a labeling process on the first category of ledger data by using a preset attribute labeling rule to obtain an attribute label corresponding to the first category of ledger data; perform a labeling process on the first category of ledger data by using a preset service labeling rule to obtain a service label corresponding to the first category of ledger data; and obtain label information of the first category of ledger data based on the attribute label corresponding to the first category of ledger data and / or the service label corresponding to the first category of ledger data.

[0084] Optionally, in a possible implementation of this embodiment, the second category of ledger data includes the entity type to which the ledger belongs, the organization to which the ledger belongs, the region to which the ledger belongs, and the ledger text description. The preset label prediction model includes a fully connected network. The prediction unit 404 is configured to input the entity type to which the ledger belongs, the organization to which the ledger belongs, the region to which the ledger belongs, and the ledger text description into the preset label prediction model; perform an encoding process on the entity type to which the ledger belongs, the organization to which the ledger belongs, the region to which the ledger belongs, and the ledger text description respectively to obtain a type encoding feature, an organization encoding feature, a region encoding feature, and a text encoding feature; perform a splicing process on the type encoding feature, the organization encoding feature, the region encoding feature, and the text encoding feature to obtain a fusion feature; input the fusion feature into the fully connected network to obtain a feature vector corresponding to the fusion feature; and obtain label information of the second category of ledger data from a preset label feature vector database based on the feature vector corresponding to the fusion feature.

[0085] Optionally, in a possible implementation of this embodiment, the preset label prediction model further includes a graph attention encoding network and a text embedding encoding network. The prediction unit 404 is configured to respectively use the graph attention encoding network to perform an encoding process on the entity type to which the ledger belongs, the organization to which the ledger belongs, and the region to which the ledger belongs to obtain the type encoding feature, the organization encoding feature, and the region encoding feature; and use the text embedding encoding network to perform an encoding process on the ledger text description to obtain a text encoding feature.

[0086] Optionally, in a possible implementation manner of this embodiment, the prediction unit 404 is configured to obtain sample ledger data and the label information corresponding to the sample ledger data; the sample ledger data includes the entity type to which the sample ledger belongs, the organization to which the sample ledger belongs, the region to which the sample ledger belongs, and the text description of the sample ledger; input the entity type to which the sample ledger belongs, the organization to which the sample ledger belongs, the region to which the sample ledger belongs, the text description of the sample ledger, and the label information corresponding to the sample ledger data into a label prediction model to be trained; respectively perform encoding processing on the entity type to which the sample ledger belongs, the organization to which the sample ledger belongs, the region to which the sample ledger belongs, and the text description of the sample ledger to obtain a type encoding feature, an organization encoding feature, a region encoding feature, and a text encoding feature; perform splicing processing on the type encoding feature, the organization encoding feature, the region encoding feature, and the text encoding feature to obtain a fusion feature; input the fusion feature into the fully connected network of the label prediction model to be trained to obtain a feature vector corresponding to the fusion feature; perform encoding processing on the label information corresponding to the sample ledger data to obtain a label feature vector; and update and train the label prediction model to be trained based on the label feature vector, the feature vector corresponding to the fusion feature, the sample ledger data, and the label information to obtain a trained label prediction model.

[0087] In this embodiment, the ledger data of the safety supervision object to be processed can be obtained by the acquisition unit, and the classification unit uses a preset classification rule to classify the ledger data to obtain first-category ledger data and second-category ledger data. The labeling unit uses a preset labeling strategy to label the first-category ledger data to obtain the label information of the first-category ledger data. The prediction unit inputs the second-category ledger data into a preset label prediction model to obtain the label information of the second-category ledger data, so that the determination unit determines the label information of the ledger data based on the label information of the first-category ledger data and the label information of the second-category ledger data. Since the label information of the ledger data can be determined by using a preset labeling strategy and a preset label prediction model respectively, and the labeling process can be carried out separately according to the type of the ledger data, the reliability and accuracy of the label information of the ledger data are improved.

[0088] In the technical solution of this application, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved, such as the user's image and attribute data, etc., all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0089] According to the embodiments of this application, this application also provides an electronic device, a readable storage medium, and a computer program product.

[0090] Figure 5 FIG. Figure 5 shows a schematic block diagram of an exemplary electronic device 500 that can be used to implement an embodiment of the present application. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present application described and / or claimed herein.

[0091] As Figure 5 shown, the electronic device 500 includes a computing unit 501 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0092] A plurality of components in the electronic device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, an optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the electronic device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0093] The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 executes the various methods and processes described above, such as the method for determining tags of ledger data. For example, in some embodiments, the method for determining tags of ledger data can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the method for determining tags of ledger data described above can be executed. Alternatively, in other embodiments, the computing unit 501 can be configured to execute the method for determining tags of ledger data by any other suitable means (e.g., by means of firmware).

[0094] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0095] The program code for implementing the methods of the present application can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.

[0096] In the context of this application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0097] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).

[0098] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of a communication network include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0099] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0100] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in the present disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solutions disclosed in the present application can be achieved, and no limitations are imposed herein.

[0101] The above specific embodiments do not constitute a limitation on the scope of protection of the present application. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present application shall be included within the scope of protection of the present application.

Claims

1. A method for determining tags of ledger data, characterized in that, The method includes: Obtaining ledger data of a safety supervision object to be processed; Using a preset classification rule to classify the ledger data to obtain first-category ledger data and second-category ledger data; Using a preset tagging strategy to tag the first-category ledger data to obtain tag information of the first-category ledger data; Inputting the second-category ledger data into a preset tag prediction model to obtain tag information of the second-category ledger data; Based on the tag information of the first-category ledger data and the tag information of the second-category ledger data, determining the tag information of the ledger data.

2. The method according to claim 1, characterized in that, The using a preset classification rule to classify the ledger data to obtain first-category ledger data and second-category ledger data includes: In response to the ledger data including not only the ledger-owned entity type, the ledger-owned organization, the ledger-owned region, and the ledger text description, determining the ledger data as first-category ledger data; In response to the ledger data only including the ledger-owned entity type, the ledger-owned organization, the ledger-owned region, and the ledger text description, determining the ledger data as second-category ledger data.

3. The method according to claim 1, characterized in that, The using a preset tagging strategy to tag the first-category ledger data to obtain tag information of the first-category ledger data includes: Using a preset attribute tag rule to tag the first-category ledger data to obtain an attribute tag corresponding to the first-category ledger data; Using a preset business tag rule to tag the first-category ledger data to obtain a business tag corresponding to the first-category ledger data; Based on the attribute tag corresponding to the first-category ledger data, and / or, the business tag corresponding to the first-category ledger data, obtaining the tag information of the first-category ledger data.

4. The method according to claim 1, characterized in that, The second-category ledger data includes the ledger-owned entity type, the ledger-owned organization, the ledger-owned region, and the ledger text description. The preset tag prediction model includes a fully connected network. The inputting the second-category ledger data into the preset tag prediction model to obtain tag information of the second-category ledger data includes: Inputting the ledger-owned entity type, the ledger-owned organization, the ledger-owned region, and the ledger text description into the preset tag prediction model; Encoding the ledger-owned entity type, the ledger-owned organization, the ledger-owned region, and the ledger text description respectively to obtain a type encoding feature, an organization encoding feature, a region encoding feature, and a text encoding feature; Performing splicing processing on the type encoding feature, the organization encoding feature, the region encoding feature, and the text encoding feature to obtain a fusion feature; Inputting the fusion feature into the fully connected network to obtain a feature vector corresponding to the fusion feature; Based on the feature vector corresponding to the fusion feature, obtaining the tag information of the second-category ledger data from a preset tag feature vector database.

5. The method according to claim 4, characterized in that, The preset label prediction model further includes a graph attention encoding network and a text embedding encoding network, which respectively perform encoding processing on the entity type to which the ledger belongs, the organization to which the ledger belongs, the region to which the ledger belongs, and the ledger text description to obtain a type encoding feature, an organization encoding feature, a region encoding feature, and a text encoding feature, including: Respectively use the graph attention encoding network to perform encoding processing on the entity type to which the ledger belongs, the organization to which the ledger belongs, and the region to which the ledger belongs to obtain the type encoding feature, the organization encoding feature, and the region encoding feature; Use the text embedding encoding network to perform encoding processing on the ledger text description to obtain a text encoding feature.

6. The method according to claim 1, characterized in that, Before inputting the second category of ledger data into the preset label prediction model to obtain the label information of the second category of ledger data, it includes: Obtain sample ledger data and the label information corresponding to the sample ledger data; the sample ledger data includes the entity type to which the sample ledger belongs, the organization to which the sample ledger belongs, the region to which the sample ledger belongs, and the sample ledger text description; Input the entity type to which the sample ledger belongs, the organization to which the sample ledger belongs, the region to which the sample ledger belongs, the sample ledger text description, and the label information corresponding to the sample ledger data into the label prediction model to be trained; Respectively perform encoding processing on the entity type to which the sample ledger belongs, the organization to which the sample ledger belongs, the region to which the sample ledger belongs, and the sample ledger text description to obtain a type encoding feature, an organization encoding feature, a region encoding feature, and a text encoding feature; Perform splicing processing on the type encoding feature, the organization encoding feature, the region encoding feature, and the text encoding feature to obtain a fusion feature; Input the fusion feature into the fully connected network of the label prediction model to be trained to obtain a feature vector corresponding to the fusion feature; Perform encoding processing on the label information corresponding to the sample ledger data to obtain a label feature vector; Based on the label feature vector, the feature vector corresponding to the fusion feature, the sample ledger data, and the label information, update and train the label prediction model to be trained to obtain a trained label prediction model.

7. A device for determining tags of ledger data, characterized in that, The device includes: An acquisition unit for acquiring the ledger data of the safety supervision object to be processed; A classification unit for classifying the ledger data using a preset classification rule to obtain first category ledger data and second category ledger data; A labeling unit for labeling the first category ledger data using a preset labeling strategy to obtain the label information of the first category ledger data; A prediction unit for inputting the second category ledger data into a preset label prediction model to obtain the label information of the second category ledger data; A determination unit for determining the label information of the ledger data based on the label information of the first category ledger data and the label information of the second category ledger data.

8. An electronic device, characterized in that, It includes: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-6.

9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-6.

10. A computer program product, characterized in that, It includes a computer program which, when executed by a processor, implements the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Method and system for determining correlation between oil chromatographic data and machine account data

    CN109214455A

  • Sample data processing method, device and system, storage medium and electronic equipment

    CN112364923A

  • Data classification method and device and storage medium

    CN113868497A

  • Graft copolymer, method for preparing the copoymer and resin composition comprising the copolymer

    KR1020210141332A

  • Methods and apparatuses for training prediction model

    US20230409929A1