Office data verification method and device, electronic equipment, storage medium and computer program product

By applying a combination of semi-supervised identification algorithm and preset rules in cloud computing products, the bureau data entities are automatically identified and verified, and the problems of low manual verification efficiency and low accuracy in the prior art are solved, and efficient and accurate automated verification management is achieved.

CN119940364APending Publication Date: 2025-05-06CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411896245.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the prior art, the local data verification of cloud computing products relies on manual configuration and approval, resulting in inefficiency and inaccurate verification.

Method used

采用基于半监督识别算法的预设分类模型和预设规则,自动识别和校验局数据实体,提高校验效率和准确率。

Benefits of technology

Automatic and efficient local data verification management is realized, which significantly improves verification efficiency and accuracy, and reduces errors in manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940364A_ABST
    Figure CN119940364A_ABST
Patent Text Reader

Abstract

The invention provides a bureau data verification method and device, electronic equipment, a storage medium and a computer program product, and relates to the technical field of data processing, and the method comprises the steps: processing a data set based on a preset classification model, and determining bureau data entities in the data set; wherein the preset classification model is determined by training a sample set determined in an entity library by using a semi-supervised recognition algorithm; and checking the bureau data entity based on a preset rule, and determining a checking result corresponding to the bureau data entity. Through the scheme in the embodiment of the invention, the verification efficiency of the bureau data and the verification accuracy of the bureau data can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a local data verification method, device, electronic device, storage medium and computer program product. Background Art

[0002] In the related technology, the process of obtaining the overall data of cloud computing products is that each product owner (PO) or product manager manually configures it through a page, and then passes it through the system approval flow, and all parties manually verify it. This process not only consumes a lot of manpower and is inefficient, but also easily causes inaccurate verification of overall data. Summary of the invention

[0003] The embodiments of the present application provide a local data verification method, device, electronic device, storage medium and computer program product, which can improve the verification efficiency and accuracy of local data.

[0004] The technical solution of this application is implemented as follows:

[0005] The present application embodiment provides a local data verification method, including:

[0006] Processing a data set based on a preset classification model to determine a local data entity in the data set; wherein the preset classification model is determined by training a sample set determined in an entity library using a semi-supervised recognition algorithm;

[0007] The local data entity is verified based on a preset rule to determine a verification result corresponding to the local data entity.

[0008] In the above scheme, the method further comprises:

[0009] Based on a semi-supervised recognition algorithm, a plurality of entities to be processed in the entity library are identified to determine the sample set in the entity library;

[0010] The initial classification model is trained based on the sample set until a predetermined training condition is reached, and the preset classification model is determined.

[0011] In the above solution, the method of identifying multiple entities to be processed in the entity library based on the semi-supervised recognition algorithm to determine the sample set in the entity library includes:

[0012] Based on the acquired first entity set, a second entity set is identified from the plurality of entities to be processed; wherein the first entity set is determined in response to a user's marking operation on the plurality of entities to be processed;

[0013] Based on the similarity between each first entity in the first entity set and each second entity in the second entity set, expanding the first entity set to determine a third entity set;

[0014] The first entity set, the third entity set and the fourth entity set are determined as the sample sets; wherein the fourth entity set is a set of entities that are not identified in the plurality of entities to be processed.

[0015] In the above solution, the first entity in the first entity set includes: a plurality of first product entities, and related entities corresponding to each of the first product entities;

[0016] The step of identifying a second entity set from among the plurality of entities to be processed based on the acquired first entity set includes:

[0017] Based on a preset model, learn a plurality of the first product entities and related entities corresponding to each of the first product entities, and output a corresponding entity set and a related entity set;

[0018] The second entity set is determined based on identifying the entity set and the related entity set in the plurality of entities to be processed.

[0019] In the above solution, the step of expanding the first entity set to determine the third entity set based on the similarity between each first entity in the first entity set and each second entity in the second entity set includes:

[0020] Determine the similarity between each of the first entities and each of the second entities based on the importance of the first keyword in the corresponding first context information and the importance of the second keyword in the corresponding second context information of each of the second entities;

[0021] Based on the similarity between each of the first entities and each of the second entities, the first entity set is expanded to determine the third entity set.

[0022] In the above solution, the determining the similarity between each first entity and each second entity based on the importance of the first keyword in each first entity in the corresponding first context information and the importance of the second keyword in each second entity in the corresponding second context information includes:

[0023] Determine a first word frequency and a first reverse document frequency of the first keyword in each of the first entities in the first context information, and determine a first feature corresponding to each of the first entities based on the first word frequency and the first reverse document frequency;

[0024] Determine a second word frequency and a second reverse document frequency of the second keyword in each of the second entities in the second context information, and determine a second feature corresponding to each of the second entities based on the second word frequency and the second reverse document frequency;

[0025] Based on the first feature and the second feature, a similarity between each of the first entities and each of the second entities is determined.

[0026] In the above solution, the initial classification model is trained based on the sample set until a predetermined training condition is reached, and the preset classification model is determined, including:

[0027] The first entity set and the third entity set are used as training inputs of the initial classification model, the fourth entity set is used as the hidden state of the initial classification model, the initial classification model is trained until a predetermined training condition is reached, and the preset classification model is determined.

[0028] In the above scheme, the method further comprises:

[0029] The preset rule is determined based on the corresponding relationship between the acquired entity and the enumeration value, and the corresponding relationship between the acquired entity and the preset regular condition.

[0030] The embodiment of the present application also provides a local data verification device, including:

[0031] A data processing unit, used to process a data set based on a preset classification model to determine a local data entity in the data set; wherein the preset classification model is determined by training a sample set determined in an entity library using a semi-supervised recognition algorithm;

[0032] The verification unit is used to verify the local data entity based on a preset rule and determine a verification result corresponding to the local data.

[0033] An embodiment of the present application further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and the processor implements the steps in the above method when executing the computer program.

[0034] An embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method are implemented.

[0035] An embodiment of the present application also provides a computer program product, including a computer program, which implements the steps in the above method when executed by a processor.

[0036] In the embodiment of the present application, a data set is processed based on a preset classification model to determine the local data entities in the data set; wherein the preset classification model is determined by training a sample set determined in an entity library using a semi-supervised recognition algorithm; and the local data entity is verified based on preset rules to determine the verification result corresponding to the local data entity. In this way, compared with the manual verification scheme in the related art, the local data is identified by a preset classification model and then verified by preset rules, which realizes automated and efficient unified verification management, thereby improving the verification efficiency of local data. In addition, since the preset classification model is determined by training a sample set determined in an entity library using a semi-supervised recognition algorithm, the preset classification model uses a large amount of unlabeled data during the training process. Semi-supervised learning can significantly improve the performance of the model under limited labeled data, resulting in a relatively high classification accuracy of the preset classification model, thereby improving the verification accuracy of local data. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 An optional flow chart of the local data verification method provided in the embodiment of the present application;

[0038] Figure 2 An optional flow chart of the local data verification method provided in the embodiment of the present application;

[0039] Figure 3 A schematic diagram of an optional effect of the local data verification method provided in an embodiment of the present application;

[0040] Figure 4 An optional flow chart of the local data verification method provided in the embodiment of the present application;

[0041] Figure 5 An optional flow chart of the local data verification method provided in the embodiment of the present application;

[0042] Figure 6 An optional flowchart of a related technical method provided in an embodiment of the present application;

[0043] Figure 7 An optional flow chart of the local data verification method provided in the embodiment of the present application;

[0044] Figure 8 An optional flow chart of the local data verification method provided in the embodiment of the present application;

[0045] Fig. 9 A schematic diagram of an optional effect of the local data verification method provided in an embodiment of the present application;

[0046] Fig.10An optional flow chart of the local data verification method provided in the embodiment of the present application;

[0047] Fig.11 An optional flow chart of the local data verification method provided in the embodiment of the present application;

[0048] Fig.12 An optional flow chart of the local data verification method provided in the embodiment of the present application;

[0049] Fig.13 A schematic diagram of the structure of a local data verification device provided in an embodiment of the present application;

[0050] Fig.14 A hardware entity schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0051] In order to make the purpose, technical solutions and advantages of the present application clearer, the technical solutions of the present application are further elaborated in detail below in conjunction with the drawings and embodiments. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.

[0052] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0053] If similar descriptions of "first / second" appear in the application documents, the following instructions are added. In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0055] In related technologies, cloud computing product data is divided into three parts: product information, specification information, and tariff information. Product information includes product name, product type, product status, etc.; specification information includes specification name, specification status, specification resource pool availability zone, etc.; tariff information includes tariff name, tariff type, tariff product name, partner information, catalog price, settlement base price, reserved assessment ratio, tax rate, etc. This information is important information for cloud computing product customer acquisition, marketing, and revenue.

[0056] At present, the product and commodity data of the cloud computing industry is mainly configured through manual pages. First, product developers analyze the competitive products in the industry, and then develop related products. Before the product is put on the shelf, the product manager needs to configure the product information through the system page, and then the product PO configures the product price and other information. To ensure the accuracy of the information configuration, manual approval is required from multiple links such as product managers, POs, operations, marketing, and department heads before the data is entered into the database. For the specific process, see Figure 1 . S01. Product & specification information. S02. Rate information. Bureau data verification is the process of each product PO or product manager manually configuring the page to obtain product & specification information and rate information. S03. Multi-role manual mutual review. Then it is approved by the system and all parties conduct manual mutual review. S04. Data processing and warehousing. After manual mutual review by multiple roles, data processing and warehousing are finally realized. This process not only consumes a lot of manpower, but also easily leads to errors in filling out information such as product name, specification name, rate name, partner information, catalog price, settlement base price, reserved assessment ratio, tax rate, etc., resulting in inaccurate billing and settlement amounts, and inaccurate verification of bureau data.

[0057] In order to solve the above technical problems, the present application embodiment provides a local data verification method, please refer to Figure 2 , is an optional flow chart of the local data verification method provided in the embodiment of the present application, which will be combined with Figure 2 The steps shown are explained:

[0058] S101. Process a data set based on a preset classification model to determine local data entities in the data set; wherein the preset classification model is determined by training a sample set determined in an entity library using a semi-supervised recognition algorithm.

[0059] In the embodiment of the present application, it is necessary to train and determine the preset classification model before identifying the local data entity. The verification device obtains the marked entities marked by the user in the entity library, and then uses the semi-supervised recognition algorithm to identify some entities in the unmarked entities in the entity library, and determines the marked entities and the identified partial entities as the sample set to train the initial classification model, and stops when the predetermined training conditions are met to determine the preset classification model. After the verification device obtains the data set, the data to be identified in the data set can be input into the preset classification model to determine the local data entity therein.

[0060] The data set may include project materials for cloud computing products, which will record product information, specification information, and price information in detail. The entity library includes multiple data entities of various types to be processed. In some other embodiments, the entity library may also be a data set. The sample set is determined in the entity library using a semi-supervised recognition algorithm, that is, the sample set is determined in the data set using a semi-supervised recognition algorithm.

[0061] Among them, the operator switch involves many concepts of "bureaus", such as exchange bureaus and tandem bureaus. Operators can be called bureaus. In the traditional sense, bureau data refers to all system data in the equipment related to system configuration, service configuration, routing organization, and network management except user data. In this application, bureau data refers to the business configuration data of cloud computing industry products. Bureau data entities may include: product information entity, specification information entity, and tariff information entity. Product information entity is a production-oriented entity. Product information entity describes the information that is necessary for the production and fulfillment of products and can be understood and recognized by the production system. Including basic product information, product attributes and other information. Commodity information entity is a sales-oriented entity that defines what products are sold and how to sell them. Conceptually, commodities are a combination of products and sales methods. Products are objects sold through commodities; and sales methods refer to a combination of contracts, sales rules, etc. The specification information entity is the smallest unit of cloud computing products for production use, and is the attribute information of the product. For example: general cloud host 2C4G; business cloud computer 4C8G, 80G system disk, 50G shared bandwidth. The fee information entity is the smallest unit of cloud computing products for sales purposes. It refers to the fees that customers pay periodically during the validity period of the transaction, such as 98 yuan per month.

[0062] S102: Verify the local data entity based on a preset rule to determine a verification result corresponding to the local data entity.

[0063] In the embodiment of the present application, it is necessary to determine the preset rules before verifying the bureau data. The verification device obtains the enumeration value of the entity input by the user and the regular conditions corresponding to other entities through the human-computer interaction device, determines the preset rules based on the obtained enumeration value and regular conditions, and then uses the preset rules to verify the identified bureau data entity to determine the verification result corresponding to the bureau data entity. If the verification result indicates that the identified bureau data entity does not comply with the preset rules, an early warning prompt can be issued for the bureau data entity.

[0064] Example: The data in the data set may include: ordering a 4-core 8G / SPECI_NAME general-purpose cloud host / PROD_NAME, if it is ordered for one year, its cost is 100 yuan / LIST_PRICE, because the settlement price with Beijing Technology Co., Ltd. / PARTNER_NAME is 80 yuan / SETTLE_PRICE, and the reserved assessment ratio is 20% / RESERVE_ASSES. The data set is identified by the preset classification model, and the identified bureau data entities can be explained by Table 1:

[0065] Serial number Entity Type Category Tags 1 Product Name PROD_NAME 2 Product Type PROD_CLASS 3 Product Status PROD_STATUS 4 Specification name SPECI_NAME 5 Price Name RATE_NAME 6 Rate Type RATE_TYPE 7 IRS Product Name TAX_BUR_NAME 8 Partner Information PARTNER_NAME 9 Catalog Price LIST_PRICE 10 Settlement base price SETTLE_PRICE 11 Reserved assessment ratio RESERVE_ASSES 12 tax rate TAX_RATE

[0066] Table 1

[0067] In the embodiment of the present application, the verification device can compare the preset rules filled in the manual page configuration with the bureau data entities identified through the above steps for consistency, and issue an early warning prompt for inconsistent bureau data entities.

[0068] This application provides a method for verifying local data. A new rule verification and matching verification mechanism is added to the module architecture to limit local data configuration content and provide error prompts, thus achieving automated and efficient unified verification management. Figure 3 , configuration interface module; mainly for product information configuration, specification information configuration and tariff information configuration. Rule matching mechanism module: mainly for information with certain rules that can be set and verified by rules, mainly for product type, product status, tariff type, catalog price, settlement base price, reserved assessment ratio, tax rate, etc. in the bureau data. The reason why this information can be verified by rules is that there are either only a few fixed enumeration values ​​or there are certain regular conditions. Algorithm matching mechanism module: mainly based on the improved semi-supervised named entity recognition algorithm to automatically identify product information, specification information, tariff information, etc. in the outbound data, and perform consistency matching verification on the identified results with the information filled in through manual page configuration. Result feedback module: mainly when the matching result representation is inconsistent, it provides inconsistent feedback, corrects the results for inconsistent bureau data, and completes information storage.

[0069] In the embodiment of the present application, a data set is processed based on a preset classification model to determine the local data entities in the data set; wherein the preset classification model is determined by training a sample set determined in an entity library using a semi-supervised recognition algorithm; and the local data entity is verified based on preset rules to determine the verification result corresponding to the local data entity. In this way, compared with the manual verification scheme in the related art, the local data is identified by a preset classification model and then verified by preset rules, which realizes automated and efficient unified verification management, thereby improving the verification efficiency of local data. In addition, since the preset classification model is determined by training a sample set determined in an entity library using a semi-supervised recognition algorithm, the preset classification model uses a large amount of unlabeled data during the training process. Semi-supervised learning can significantly improve the performance of the model under limited labeled data, resulting in a relatively high classification accuracy of the preset classification model, thereby improving the verification accuracy of local data.

[0070] See also Figure 4 , which is an optional flow chart of the local data verification method provided in the embodiment of the present application, will be described in combination with the steps:

[0071] S201. Identify a plurality of entities to be processed in an entity library based on a semi-supervised recognition algorithm, and determine the sample set in the entity library.

[0072] In the embodiment of the present application, the entity library includes multiple entities to be processed of various types. The verification device can mark a first entity set among the multiple entities to be processed in response to the user's operation, and determine the unmarked entities, and use a semi-supervised recognition algorithm to identify a large number of entities that meet the matching requirements based on the marked first entity set combined with a large number of unmarked entities as a sample set.

[0073] In the embodiment of the present application, the verification device can use a semi-supervised recognition algorithm to identify the second entity set based on the marked first entity set combined with a large number of unmarked entities, and then expand the first entity to determine the third entity based on the similarity between the entities in the second entity set and the entities in the first entity set. The first entity set and the second entity set are determined as the sample set.

[0074] S202: Train the initial classification model based on the sample set until a predetermined training condition is reached, and then determine the preset classification model.

[0075] In the embodiment of the present application, the verification device can train the initial classification model in combination with the sample set and the unidentified entities in the entity set until a predetermined number of training times is reached or the function converges, and then determine the preset classification model.

[0076] The initial classification model may include a hierarchical conditional random field (HCRF) model. The verification device may train the HCRF model in combination with the sample set and the unidentified entities in the entity set to train a classifier. In other embodiments, the initial classification model may also be other types of models. In the embodiments of the present application, there is no specific limitation on the initial classification model.

[0077] In the embodiment of the present application, a sample set in the entity library is determined based on a semi-supervised recognition algorithm for multiple entities to be processed in the entity library. The initial classification model is trained based on the sample set until the predetermined training condition is reached, and the preset classification model is determined. In this way, since the preset classification model is determined by training the sample set determined in the entity library using the semi-supervised recognition algorithm, the preset classification model uses a large amount of unlabeled data during the training process. Semi-supervised learning can significantly improve the performance of the model under limited labeled data, resulting in a relatively high classification accuracy of the preset classification model, thereby improving the verification accuracy of the local data.

[0078] See also Figure 5 , is an optional flow chart of the local data verification method provided in the embodiment of the present application, Figure 4 The illustrated S201 can also be implemented through S301 to S303, which will be described in conjunction with the steps:

[0079] S301. Identify a second entity set from the plurality of entities to be processed based on the acquired first entity set; wherein the first entity set is determined in response to a user's marking operation on the plurality of entities to be processed.

[0080] In the embodiment of the present application, the verification device can mark a first entity set including multiple first entities in the entity set in response to the marking operation of the user through a human-computer interaction device, and then use each first entity to match and identify multiple second entities in the unmarked entities in the multiple entities to be processed to form a second entity set.

[0081] S302: Expand the first entity set to determine a third entity set based on the similarity between each first entity in the first entity set and each second entity in the second entity set.

[0082] In the embodiment of the present application, the verification device compares the similarity between each second entity and each first entity to determine the similarity between each second entity and each first entity. The second entities with similarity greater than a preset threshold are added to the first entity set to determine the third entity set.

[0083] S303: Determine the first entity set, the third entity set, and the fourth entity set as the sample set; wherein the fourth entity set is a set of entities that are not identified in the plurality of entities to be processed.

[0084] In the embodiment of the present application, the verification device can use the first entity set and the third entity set as input, and use the fourth entity set as a hidden state to train the initial classification model until the predetermined training condition is reached, and then stop to determine the preset classification model. The fourth entity set is a set of entities that are not identified in the plurality of entities to be processed.

[0085] In an embodiment of the present application, a second entity set is identified among multiple entities to be processed based on the acquired first entity set; wherein the first entity set is determined in response to a user's marking operation on multiple entities to be processed. Based on the similarity between each first entity in the first entity set and each second entity in the second entity set, the first entity set is expanded to determine a third entity set. The first entity set, the third entity set, and the fourth entity set are determined as sample sets; wherein the fourth entity set is a set of entities that have not been identified among the multiple entities to be processed. In an embodiment of the present application, a second entity set is identified among multiple entities to be processed using the first entity set, and then the first entity set is expanded using the second entity to determine the third entity set. A sufficient number of samples can be obtained, and then the initial classification model can be fully trained, and a preset classification model with better performance can be obtained to achieve relatively high accuracy classification of local data.

[0086] See also Figure 6 , is an optional flow chart of the local data verification method provided in the embodiment of the present application, Figure 5 The illustrated S301 can also be implemented through S401 to S402, which will be described in conjunction with the steps:

[0087] S401. Learn multiple first product entities and related entities corresponding to each of the first product entities based on a preset model, and output corresponding entity sets and related entity sets.

[0088] In the embodiment of the present application, the first entity in the first entity set includes: multiple first product entities, and related entities corresponding to each of the first product entities. The related entities corresponding to each of the first product entities may include: a specification information entity corresponding to the first product entity, and a tariff information entity corresponding to the first product entity. The verification device can use a preset model to learn each first product entity in the first entity set, as well as a specification information entity corresponding to each first product entity, and a tariff information entity corresponding to each first product entity. Then, the entity set and the related entity set are output through the learned preset model.

[0089] Among them, the preset model may include a generative pre-trained transformer (GPT) model. In other embodiments, the preset model may also be other types of models, which are not specifically limited in the embodiments of the present application. GPT is a natural language processing model architecture based on deep learning. It is pre-trained through large-scale text data and can understand and generate human language.

[0090] The first product entity set may include: a brand dictionary, a series dictionary, such as a general-purpose cloud host, a computing cloud host, a memory cloud host, etc.

[0091] The related entity sets corresponding to the first product entity set may include product specifications and models (combinations of numbers and letters), such as 2 cores 4G, 2 cores 8G, standard version with two copies 32GB, etc., as well as combination relationships, where models and versions are combined into specifications, and specifications + billing cycles are combined into tariffs, such as standard package (16 cores 64G), monthly package (standard version with two copies 32GB), etc.

[0092] S402: Determine the second entity set based on identifying the entity set and the related entity set in the plurality of entities to be processed.

[0093] In an embodiment of the present application, the verification device can match and identify unlabeled entities in multiple entities to be processed based on the entity set output by the preset model and the related entity set, and determine the second entity set.

[0094] Exemplarily, the identified second entity set can be described by Table 2:

[0095]

[0096]

[0097] Table 2

[0098] In an embodiment of the present application, a plurality of first product entities and related entities corresponding to each first product entity are learned based on a preset model, and the corresponding entity set and related entity set are output. A second entity set is determined based on the entity set and the related entity set in a plurality of entities to be processed. In this way, a large number of second entities matching the output entity set and the related entity set can be extracted from unidentified entities using the entity set and the related entity set output by the model to form a second entity set. This solution can effectively expand the samples trained for the initial classification model, so that the initial classification model is fully trained, and the classification accuracy of the preset classification model is improved.

[0099] See also Figure 7 , is an optional flow chart of the local data verification method provided in the embodiment of the present application, Figure 5 The illustrated S302 can also be implemented through S501 to S502, which will be described in conjunction with the steps:

[0100] S501: Determine the similarity between each of the first entities and each of the second entities based on the importance of the first keyword in each of the first entities in the corresponding first context information and the importance of the second keyword in each of the second entities in the corresponding second context information.

[0101] In an embodiment of the present application, the first entity may include a corresponding first keyword and the first context information corresponding to the keyword, and the first context information is text information used to explain and supplement the first keyword. The verification device determines the feature value of the corresponding first entity based on the importance of the first keyword in the corresponding first context information in each first entity. Similarly, the feature value of the corresponding second entity is determined based on the importance of the second keyword in the corresponding second context information in each second entity. The similarity between each second entity and each first entity can be calculated by the feature value of each second entity and the feature value of each first entity.

[0102] In an embodiment of the present application, when determining the feature value, the verification device may consider the keyword attributes used to characterize the importance of the entity, which may include the word frequency and reverse document frequency of the entity in the corresponding context information. In some other embodiments, when determining the feature value corresponding to the entity, other attributes corresponding to the entity may also be considered, which may include attributes such as relevance, flow, and difficulty.

[0103] S502: Based on the similarity between each of the first entities and each of the second entities, expand the first entity set and determine the third entity set.

[0104] In the embodiment of the present application, the verification device traverses each second entity, and if the similarity between a second entity and any first entity is greater than a preset threshold, the second entity is added to the first entity set until the traversal is completed and the third entity set is determined.

[0105] In the embodiment of the present application, the similarity between each first entity and each second entity is determined based on the importance of the first keyword in each first entity in the corresponding first context information, and the importance of the second keyword in each second entity in the corresponding second context information. Based on the similarity between each first entity and each second entity, the first entity set is expanded and the third entity set is determined. In this way, by finding the second entity similar to the first entity to expand the first entity set, enough samples can be obtained so that the initial classification model can be effectively trained, thereby improving the classification accuracy of the preset classification model.

[0106] See also Figure 8 , is an optional flow chart of the local data verification method provided in the embodiment of the present application, Figure 7 The illustrated S501 can also be implemented through S601 to S603, which will be described in conjunction with the steps:

[0107] S601: Determine a first word frequency and a first reverse document frequency of the first keyword in each of the first entities in the first context information, and determine a first feature corresponding to each of the first entities based on the first word frequency and the first reverse document frequency.

[0108] In the embodiment of the present application, the verification device can calculate and determine the first word frequency and the first reverse document frequency of the corresponding keyword in the first context information for each first entity in the first entity set, and then combine the first word frequency and the first reverse document frequency to determine the first feature corresponding to each first entity.

[0109] The first word frequency and the first reverse document frequency are both used to characterize the importance of the first keyword in the first context information corresponding to the first entity.

[0110] Exemplarily, the verification device may calculate the first word frequency TF corresponding to the first keyword by using formula (1).

[0111]

[0112] Among them, n w,C is the number of times the first keyword w appears in the first context information C, is the normalization factor.

[0113] Exemplarily, the verification device may calculate the first inverse document frequency IDF corresponding to the first keyword by using formula (2).

[0114]

[0115] Wherein, N is the total number of first context information in the first entity, n is the number of first keyword w included in the first context information, and n w,a is the number of times the first keyword w appears in category A, is the number of times the first keyword w appears in the first entity of other categories.

[0116] For example, the verification device can calculate the first feature v corresponding to the first entity through formula (3).

[0117]

[0118] S602: Determine a second word frequency and a second reverse document frequency of the second keyword in each of the second entities in the second context information, and determine a second feature corresponding to each of the second entities based on the second word frequency and the second reverse document frequency.

[0119] In the embodiment of the present application, the verification device can calculate the second word frequency and the second reverse document frequency of the second keyword in the corresponding second context information for each second entity in the second entity set, and then combine the second word frequency and the second reverse document frequency to determine the second feature corresponding to each second entity.

[0120] The second word frequency and the second inverse document frequency are both used to characterize the importance of the second keyword in the corresponding second entity in the corresponding second context information.

[0121] In the embodiment of the present application, the verification device can refer to the above formulas (1) to (3) to determine the second word frequency, the second reverse document frequency and the second feature corresponding to each second entity.

[0122] S603: Determine the similarity between each of the first entities and each of the second entities based on the first feature and the second feature.

[0123] In the embodiment of the present application, after determining the first feature corresponding to the first entity and the second feature corresponding to the second entity, the similarity between each second feature and each first feature can be calculated using a similarity calculation method for each second feature.

[0124] For example, the similarity can be calculated by Tanimoto coefficient. The Tanimoto coefficient calculation formula (4) is:

[0125]

[0126] Among them, v s is the first feature, v c is the second feature, and k is the number of corresponding features.

[0127] In the embodiment of the present application, the first word frequency and the first reverse document frequency of the first keyword in each first entity in the first context information are determined, and based on the first word frequency and the first reverse document frequency, the first feature corresponding to each first entity is determined; the second word frequency and the second reverse document frequency of the second keyword in each second entity in the second context information are determined, and based on the second word frequency and the second reverse document frequency, the second feature corresponding to each second entity is determined; based on the first feature and the second feature, the similarity between each first entity and each second entity is determined. In this way, by determining the corresponding entity features through word frequency and reverse document frequency, the features of the entity are considered from two directions respectively, and the consideration is more comprehensive, so that the features of the determined entity are more accurate.

[0128] In the embodiments of the present application, local data entities refer to entities with specific meanings in the recognized text, such as names of people, places, organizations, proper nouns, etc. However, for the named entities of product information in the cloud computing industry, due to their variable structural characteristics, fuzzy boundaries, and lack of labeled data; therefore, the structural characteristics of product information make it more difficult to identify product named entities than general named entities. General named entity recognition technology cannot effectively identify product entities. The present application provides an improved three-layer semi-supervised named entity recognition algorithm that can better complete the named entity recognition task of product information for the cloud computing industry. Based on the improved three-layer semi-supervised named entity recognition algorithm, the local data entities listed in are identified. The overall framework of the algorithm is as follows: Fig. 9 As shown:

[0129] First layer: Candidate set selection For named entities of product information in the cloud computing industry, the structural characteristics of product information are not obvious, but the content is stable, and the candidate set can be screened by establishing a dictionary. For product specification information, it is difficult to establish a dictionary because new specifications are often launched, but specifications are generally a combination of numbers, models and letters, and the specification information can be matched through the specification structure. For product tariff information, the specifications and billing cycles are combined in a certain relationship.

[0130] The second layer: expansion of the positive example set. The candidate set selected based on a specific strategy needs to be further identified as a named entity of product information, a named entity of other categories, or just an ordinary word. Through an improved context similarity calculation method, the context of the candidate set and the context of the positive example set can be similarly calculated to find candidate words similar to the positive example and add them to the positive example, thereby obtaining sufficient learning labeled data. At the same time, according to the increase in the size of the positive example set, the rule base and dictionary are updated to improve the accuracy and coverage of the candidate set selection.

[0131] The third layer: CRF model fused with attention mechanism classifier. The contextual similarity comparison method can filter out candidate words with high similarity, but some entities with low similarity cannot be recognized. In the third layer, the positive example annotations obtained through the first and second layers are used as training inputs, and the unrecognized entities are used as hidden states. A classifier is trained through the HCRF model to output the final cloud computing product named entity results.

[0132] See also Fig.10 , is an optional flow chart of the local data verification method provided in the embodiment of the present application, Figure 5 The illustrated S202 can also be implemented by S701, which will be described in conjunction with the steps:

[0133] S701. Use the first entity set and the third entity set as training inputs of the initial classification model, use the fourth entity set as a hidden state of the initial classification model, train the initial classification model until a predetermined training condition is met, and determine the preset classification model.

[0134] In an embodiment of the present application, the initial classification model may be an HCRF model. The verification device may use the first entity set and the third entity set as training inputs of the initial classification model, use the fourth entity set as a hidden state of the initial classification model, train the initial classification model until a predetermined training condition is reached, and stop to determine the preset classification model.

[0135] Among them, the HCRF model is a machine learning model for sequence labeling tasks, which can learn the mapping relationship between input sequences and output labels. Assume that the entity to be identified is X = (x1, x2, ..., x L ), example X = (this is a cloud host), the corresponding training sample is Y = (y1, y2, ..., y L ), example Y = (this / O is / O a / O cloud / B host / M machine / E). Given the first entity set T = {(X1, Y1), (X2, Y2), ..., (X T ,Y T )} and the third entity set T'={(X1,Z1),(X2,Z2),...,(X Y ,Z T )}, since the third entity set only classifies the parts with high similarity, T' may be missing. If not, then If there is a missing It is the hidden state.

[0136] Assume that the fourth entity set in the training sample is h = {h1,h2,...,h M}, then the conditional probability formula (5) of the label sequence is:

[0137]

[0138] Among them, y is the marking result, h is the hidden state, x is the character to be marked, and λ is the weight. y(y,h,x;λ) is the potential function: in, It is a common feature parameter between the third entity set and the fourth entity set.

[0139] In an embodiment of the present application, the first entity set and the third entity set are used as training inputs of the initial classification model, and the fourth entity set is used as the hidden state of the initial classification model. The initial classification model is trained until the predetermined training condition is reached and the preset classification model is determined. Since the entity set for training the initial classification model of the present application is determined based on a semi-supervised recognition algorithm, a sufficient number of training samples can be determined from a limited entity set to ensure the adequacy of the training of the initial classification model, thereby improving the classification performance of the preset classification model.

[0140] See also Fig.11 , which is an optional flow chart of the local data verification method provided in the embodiment of the present application, will be described in combination with the steps:

[0141] S801: Determine the preset rule based on the corresponding relationship between the acquired entity and the enumeration value, and the corresponding relationship between the acquired entity and the preset regular condition.

[0142] In the embodiment of the present application, the verification device can obtain the enumeration value configured by the user for each entity based on the human-computer interaction device, and then obtain the corresponding relationship between the entity and the enumeration value. The verification device can also obtain the preset regularity condition configured by the user for the entity based on the human-computer interaction device, and then obtain the corresponding relationship between the entity and the corresponding preset regularity condition. Based on the obtained corresponding relationship between the entity and the enumeration value, and the obtained corresponding relationship between the entity and the preset regularity condition, the preset rule is determined.

[0143] Among them, entities with fixed enumeration values ​​are: product type (IaaS, PaaS, SaaS), product status (on the shelf, public beta, internal beta, off the shelf), rate type (annual package, monthly package, annual one-time, one-time, call bill). Entities with regular conditions are: catalog price, settlement base price, reserved assessment ratio, tax rate, among which catalog price>settlement base price, reserved assessment ratio<40%, tax rate<=6%.

[0144] In the embodiment of the present application, the verification device can add a rule limitation mechanism when filling in the fields of product type, product status, tariff type, catalog price, settlement base price, reserved assessment ratio, and tax rate based on the regular conditions of the above fields. The specific mechanism is as follows: the product type can only be one of IaaS, PaaS, and SaaS, the product status can only be selected from listing, public beta, internal beta, and delisting, and the tariff type can only be selected from annual package, monthly package, annual package one-time, one-time, and call bill. If the catalog price < settlement base price or reserved assessment ratio > 40% or the tax rate > 6%, an error message will be prompted.

[0145] See also Fig.12 , which is an optional flow chart of the local data verification method provided in the embodiment of the present application, will be described in combination with the steps:

[0146] S11. Obtain product & specification information.

[0147] S12. Obtaining tariff information.

[0148] In the embodiment of the present application, enumeration values ​​of product information entity and specification information entity are obtained, and fixed enumeration values ​​are set: product type, product status, and tariff type. And the regular conditions corresponding to the settings in the tariff information are exemplarily included: settlement base price < catalog price, reserved assessment ratio < 40%.

[0149] S13. Data processing is stored in a temporary database.

[0150] In the embodiment of the present application, corresponding preset rules are formed based on the obtained enumeration values ​​and regular conditions, which are used to verify the subsequently identified local data entities.

[0151] S14. Based on the improved semi-supervised named entity recognition algorithm fusion, product information, specification information, and price information are extracted.

[0152] In an embodiment of the present application, product project review documents and procurement contracts can be scanned, and the scanned text information can be fused through an improved semi-supervised named entity recognition algorithm to extract product information, specification information, tariff information and other local data entities.

[0153] S15. Automated matching verification.

[0154] In the embodiment of the present application, an automatic matching check is performed on the identified local data entity based on a determined preset rule.

[0155] S16: Check whether the corresponding information matches.

[0156] In the embodiment of the present application, if the bureau data entity matches the corresponding preset rule, the bureau data is processed into the data warehouse, and the subsequent bureau data billing takes effect. If the bureau data entity does not match the corresponding preset rule, an error prompt is given for the bureau data entity, so that the user can manually correct the bureau data entity.

[0157] See also Fig.13 , is a structural diagram of the local data verification device provided in an embodiment of the present application.

[0158] The embodiment of the present application also provides a local data verification device 800, including: a data processing unit 801 and a verification unit 802.

[0159] The data processing unit 801 is used to process the data set based on a preset classification model to determine the local data entity in the data set; wherein the preset classification model is determined by training a sample set determined in the entity library using a semi-supervised recognition algorithm;

[0160] The verification unit 802 is used to verify the local data entity based on a preset rule and determine a verification result corresponding to the local data entity.

[0161] In the embodiment of the present application, the data processing unit 801 in the local data verification device 800 is used to identify multiple entities to be processed in the entity library based on a semi-supervised recognition algorithm, and determine the sample set in the entity library;

[0162] The initial classification model is trained based on the sample set until a predetermined training condition is reached, and the preset classification model is determined.

[0163] In the embodiment of the present application, the data processing unit 801 in the local data verification device 800 is used to identify a second entity set from the plurality of entities to be processed based on the acquired first entity set; wherein the first entity set is determined in response to a user's marking operation on the plurality of entities to be processed;

[0164] Based on the similarity between each first entity in the first entity set and each second entity in the second entity set, expanding the first entity set to determine a third entity set;

[0165] The first entity set, the third entity set and the fourth entity set are determined as the sample sets; wherein the fourth entity set is a set of entities that are not identified in the plurality of entities to be processed.

[0166] In the embodiment of the present application, the first entity in the first entity set includes: a plurality of first product entities and related entities corresponding to each of the first product entities; the data processing unit 801 in the local data verification device 800 is used to learn the plurality of first product entities and related entities corresponding to each of the first product entities based on a preset model, and output the corresponding entity set and the related entity set;

[0167] The second entity set is determined based on identifying the entity set and the related entity set in the plurality of entities to be processed.

[0168] The data processing unit 801 in the local data verification device 800 is used to determine the similarity between each of the first entities and each of the second entities based on the importance of the first keyword in each of the first entities in the corresponding first context information and the importance of the second keyword in each of the second entities in the corresponding second context information;

[0169] Based on the similarity between each of the first entities and each of the second entities, the first entity set is expanded to determine the third entity set.

[0170] The data processing unit 801 in the local data verification device 800 is used to determine the first word frequency and the first reverse document frequency of the first keyword in each of the first entities in the first context information, and determine the first feature corresponding to each of the first entities based on the first word frequency and the first reverse document frequency;

[0171] Determine a second word frequency and a second reverse document frequency of the second keyword in each of the second entities in the second context information, and determine a second feature corresponding to each of the second entities based on the second word frequency and the second reverse document frequency;

[0172] Based on the first feature and the second feature, a similarity between each of the first entities and each of the second entities is determined.

[0173] The data processing unit 801 in the local data verification device 800 is used to use the first entity set and the third entity set as training inputs of the initial classification model, use the fourth entity set as the hidden state of the initial classification model, train the initial classification model until a predetermined training condition is reached, and stop to determine the preset classification model.

[0174] The data processing unit 801 in the local data verification device 800 is used to determine the preset rule based on the corresponding relationship between the acquired entity and the enumeration value, and the corresponding relationship between the acquired entity and the preset regular condition.

[0175] It should be noted that in the embodiments of the present application, if the above-mentioned method for processing item information is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application can be essentially or partly reflected in the form of a software product that contributes to the relevant technology. The computer software product is stored in a storage medium and includes several instructions for enabling an item information processing device (which can be a personal computer, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0176] Correspondingly, an embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps in the method on the local data verification device side are implemented.

[0177] It should be noted here that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.

[0178] It should be noted that Fig.14 A hardware entity diagram of an electronic device provided in an embodiment of the present application, such as Fig.14 As shown, an embodiment of the present application provides an electronic device 900, including a memory 902 and a processor 901, wherein the memory 902 stores a computer program that can be run on the processor 901, and the processor 901 implements the steps in the above method when executing the program, wherein;

[0179] The processor 901 generally controls the overall operation of the electronic device 900 .

[0180] The memory 902 is configured to store instructions and applications executable by the processor 901, and can also cache data to be processed or processed by the processor 901 and various modules in the electronic device 900 (for example, image data, audio data, voice communication data, and video communication data), which can be implemented through flash memory (FLASH) or random access memory (Random Access Memory, RAM).

[0181] Correspondingly, an embodiment of the present application further provides a computer program product, including a computer program, which can be executed by the processor 901 of the electronic device 900 to complete the steps in the method on the side of the local data verification device 800.

[0182] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the size of the sequence number of the above-mentioned processes does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned sequence numbers of the embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.

[0183] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0184] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be electrical, mechanical or other forms.

[0185] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0186] In addition, all functional units in the embodiments of the present application may be integrated into one processing unit, or each unit may be a separate unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0187] A person skilled in the art can understand that all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: a mobile storage device, a read-only memory (ROM), a magnetic disk or an optical disk, and other media that can store program codes.

[0188] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application can be essentially or partly embodied in the form of a software product that contributes to the relevant technology. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a mobile storage device, a ROM, a magnetic disk, or an optical disk.

[0189] The above is only an implementation method of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application.

Claims

1. A local data verification method, characterized in that: include: Processing a data set based on a preset classification model to determine a local data entity in the data set; wherein the preset classification model is determined by training a sample set determined in an entity library using a semi-supervised recognition algorithm; The local data entity is verified based on a preset rule to determine a verification result corresponding to the local data entity.

2. The local data verification method according to claim 1, characterized in that: The method further comprises: Based on a semi-supervised recognition algorithm, a plurality of entities to be processed in the entity library are identified to determine the sample set in the entity library; The initial classification model is trained based on the sample set until a predetermined training condition is reached, and the preset classification model is determined.

3. The local data verification method according to claim 2, characterized in that: The identifying multiple entities to be processed in the entity library based on the semi-supervised recognition algorithm to determine the sample set in the entity library includes: Based on the acquired first entity set, a second entity set is identified from the plurality of entities to be processed; wherein the first entity set is determined in response to a user's marking operation on the plurality of entities to be processed; Based on the similarity between each first entity in the first entity set and each second entity in the second entity set, expanding the first entity set to determine a third entity set; The first entity set, the third entity set and the fourth entity set are determined as the sample sets; wherein the fourth entity set is a set of entities that are not identified in the plurality of entities to be processed.

4. The local data verification method according to claim 3, characterized in that: The first entity in the first entity set includes: a plurality of first product entities and related entities corresponding to each of the first product entities; The step of identifying a second entity set from among the plurality of entities to be processed based on the acquired first entity set includes: Based on a preset model, learn a plurality of the first product entities and related entities corresponding to each of the first product entities, and output a corresponding entity set and a related entity set; The second entity set is determined based on identifying the entity set and the related entity set in the plurality of entities to be processed.

5. The local data verification method according to claim 3, characterized in that: The method of expanding the first entity set to determine a third entity set based on the similarity between each first entity in the first entity set and each second entity in the second entity set includes: Determine the similarity between each of the first entities and each of the second entities based on the importance of the first keyword in the corresponding first context information and the importance of the second keyword in the corresponding second context information of each of the second entities; Based on the similarity between each of the first entities and each of the second entities, the first entity set is expanded to determine the third entity set.

6. The local data verification method according to claim 5, characterized in that: The determining the similarity between each of the first entities and each of the second entities based on the importance of the first keyword in the corresponding first context information and the importance of the second keyword in the corresponding second context information of each of the second entities includes: Determine a first word frequency and a first reverse document frequency of the first keyword in each of the first entities in the first context information, and determine a first feature corresponding to each of the first entities based on the first word frequency and the first reverse document frequency; Determine a second word frequency and a second reverse document frequency of the second keyword in each of the second entities in the second context information, and determine a second feature corresponding to each of the second entities based on the second word frequency and the second reverse document frequency; Based on the first feature and the second feature, a similarity between each of the first entities and each of the second entities is determined.

7. The local data verification method according to claim 3, characterized in that: The initial classification model is trained based on the sample set until a predetermined training condition is reached and then the preset classification model is determined, including: The first entity set and the third entity set are used as training inputs of the initial classification model, the fourth entity set is used as the hidden state of the initial classification model, the initial classification model is trained until a predetermined training condition is reached, and the preset classification model is determined.

8. The local data verification method according to any one of claims 1 to 7, characterized in that: The method further comprises: The preset rule is determined based on the corresponding relationship between the acquired entity and the enumeration value, and the corresponding relationship between the acquired entity and the preset regular condition.

9. A local data verification device, characterized in that: include: A data processing unit, configured to process a data set based on a preset classification model to determine a local data entity in the data set; wherein the preset classification model is determined by training a sample set determined in an entity library using a semi-supervised recognition algorithm; The verification unit is used to verify the local data entity based on a preset rule and determine a verification result corresponding to the local data entity.

10. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program executable on the processor, and the processor implements the steps of the method according to any one of claims 1 to 8 when executing the computer program.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps in the method according to any one of claims 1 to 8 are implemented.

12. A computer program product, comprising a computer program, characterized in that The computer program implements the steps of the method according to any one of claims 1 to 8 when executed by a processor.