Cross-department device naming consistency evaluation method and system
By constructing a semantic tag matrix and using hash signatures to assess device naming consistency, the problem of inconsistent device naming across departments was solved, enabling efficient data integration and intelligent management while protecting sensitive data.
Patent Information
- Application Number
- CN202511430658.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-09
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-10-09
AI Technical Summary
Inconsistent naming of equipment across departments leads to difficulties in data sharing and integration, and existing methods rely on manual evaluation, which is inefficient and costly.
A semantic tag matrix is constructed using a tag concatenation vectorization method. The similarity between devices is calculated by combining cosine similarity and Jaccard similarity. A hash signature is generated using the standard locality-sensitive hash algorithm and privacy protection mechanism. Consistency assessment is ensured by correcting the similarity.
It improves the efficiency of equipment cataloging data integration, reduces reliance on manual labor and the difficulty of cross-departmental collaboration, and ensures consistency in equipment naming and protection of sensitive data.
Smart Images

Figure CN120912154A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of intelligent supervision, in particular to a cross-department equipment naming consistency evaluation method and system. BACKGROUND
[0002] Equipment cataloging is a basic and preparatory link of equipment data resource construction and management. Cataloging data is the core and key data for realizing equipment data governance and utilization, and the data quality thereof is directly related to the credibility and effectiveness of equipment data resource utilization. When different departments conduct equipment cataloging, experts assign specific equipment with a variety of names based on a benchmark name and in combination with equipment attributes. If the degree of communication and cooperation between experts is low, the problem of inconsistent naming of the same equipment variety across departments will occur. For example, a certain equipment is labeled as "reconnaissance unmanned aerial vehicle" or "airborne surveillance platform" in different data systems, but it is essentially the same type of equipment. In the cross-department scenario, it is also an important requirement to ensure the consistency of equipment cataloging data while preventing sensitive data leakage during data sharing and integration.
[0003] Currently, the solution to cross-department equipment naming consistency evaluation mainly relies on expert knowledge, which is high in labor cost, time-consuming and low in efficiency, and seriously restricts the construction quality of equipment cataloging data. SUMMARY
[0004] The purpose of the present application is to provide a cross-department equipment naming consistency evaluation method and system, which can effectively solve the problems of strong manual dependence, low efficiency and great difficulty in cross-department collaboration in traditional methods, and ensure efficient integration and intelligent management of equipment cataloging data.
[0005] To achieve the above-mentioned purpose, the present application provides the following solutions: In a first aspect, the present application provides a cross-department equipment naming consistency evaluation method, which comprises: constructing a semantic label matrix by using a label splicing vectorization method according to multi-label annotation data of equipment samples; determining the average similarity between different equipment samples by using cosine similarity and Jaccard similarity according to the semantic label matrix; determining the hash signature between different equipment samples by using a standard local sensitive hashing algorithm and a privacy protection mechanism according to the semantic label matrix; the privacy protection mechanism is to randomly flip the hash signature bit according to the privacy level of the equipment; determining the similarity of the hash signature between different equipment samples, and correcting the similarity by using the privacy level.
[0006] Optionally, the construction of the semantic label matrix by using the label splicing vectorization method according to the multi-label annotation data of the equipment samples specifically comprises: using formula constructing multi-annotator label concatenation vector ; using constructing semantic label matrix ; wherein, is the label vector of the device sample annotated by the annotator , is the number of annotators, is the category of the device sample, is the number of device samples, represents the concatenation operation of the vector, and the superscript T represents the transpose.
[0007] Optionally, the label concatenation vectorization method is used to construct the semantic label matrix according to the multi-label annotation data of the device sample, and the method further comprises the following steps: detecting errors in the multi-label annotation data.
[0008] Optionally, the average similarity between different device samples is determined according to the semantic label matrix by using cosine similarity and Jaccard similarity, and the method specifically comprises the following steps: determining the cosine similarity and the Jaccard similarity between two device samples according to the semantic label matrix; determining the average similarity by using a weighted average method according to the cosine similarity and the Jaccard similarity between the two device samples.
[0009] Optionally, the average similarity is determined by using a weighted average method according to the cosine similarity and the Jaccard similarity between the two device samples, and the method specifically comprises the following steps: using formula to determine the average similarity between the device sample and the device sample ; wherein, is a weighted coefficient, is the cosine similarity between the device sample and the device sample , is the Jaccard similarity between the device sample and the device sample .
[0010] Optionally, the hash signature between different device samples is determined according to the semantic label matrix by using a standard local sensitive hashing algorithm and a privacy protection mechanism, and the method specifically comprises the following steps: using formula determining a hash signature of a device sample i ; according to a privacy protection level generating a vector of randomly flipping hash bits in the hash signature ; using a formula determining the flipped hash signature ; wherein, is a hash bit generated after k times of random projection of the device sample i, denotes a bitwise XOR operation, and each element of the random vector is generated by an independent Bernoulli random variable with a probability of the privacy protection level .
[0011] Optionally, the similarity of the hash signatures between different device samples is determined, and the similarity is corrected using a privacy level, specifically including: using a formula determining the similarity of the hash signatures between different device samples ; using a formula determining the similarity of the corrected hash signatures ; wherein, and denote the privacy-protected hash values of two device samples or annotators, is the number of bits of the hash signature, is an indicator function, and is 1 if , otherwise 0, which is used to count the matching of the two hash signatures at each hash bit, is the privacy protection level, indicating how many hash bits are randomly flipped.
[0012] In a second aspect, the present application provides a cross-department device naming consistency evaluation system, which comprises: a label semantic vectorization module, configured to construct a semantic label matrix by using a label splicing vectorization method according to multi-label annotation data of device samples; a consistency evaluation module, configured to determine the average similarity between different device samples by using cosine similarity and Jaccard similarity according to the semantic label matrix; a privacy protection hash matching module, configured to determine the hash signatures between different device samples by using a standard local sensitive hash algorithm and a privacy protection mechanism according to the semantic label matrix; the privacy protection mechanism is to randomly flip the hash signature bits according to the privacy level of the device; The consistency evaluation and privacy protection similarity correction module is configured to determine the similarity of hash signatures between different device samples and correct the similarity by using a privacy level.
[0013] According to the specific embodiments provided in the application, the application has the following technical effects: The application provides a cross-department device naming consistency evaluation method and system. According to the multi-label annotation data of the device samples, a label splicing vectorization method is used to construct a semantic label matrix, and the original device label data is converted into a high-dimensional semantic vector to support subsequent consistency evaluation. On the basis of the hash signature of the standard locality-sensitive hashing (LSH) algorithm, efficient approximate similarity calculation of the high-dimensional semantic vector is realized, and a random disturbance mechanism is introduced. The signature bits are randomly flipped according to the set privacy level, so as to desensitize and encrypt the naming data. In view of the error caused by the disturbance, a corresponding numerical correction model is constructed to repair the deviation of the disturbed similarity, so as to restore the original semantic consistency distribution. Through the implementation of this scheme, the sensitive device data protection can be realized while improving the data integration efficiency when performing cross-department data consistency evaluation, greatly reducing the data ambiguity and misjudgment risk caused by inconsistent naming. The problems of strong manual dependence, low efficiency and great difficulty in cross-department cooperation in the traditional method can be effectively solved, and strong support is provided for efficient integration and intelligent management of device cataloging data. BRIEF DESCRIPTION OF DRAWINGS
[0014] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the drawings needed in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description only constitute some embodiments of the application, and for those skilled in the art, other drawings can be obtained based on these drawings without creative labor.
[0015] Figure 1 A cross-department device naming consistency evaluation method flowchart in an embodiment of the application; Figure 2 A cross-department device naming consistency evaluation method flowchart in an embodiment of the application; Figure 3 A cross-department device naming consistency evaluation system structure diagram in an embodiment of the application. DETAILED DESCRIPTION
[0016] With reference to the drawings, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all the other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0017] The above-mentioned purposes, features and advantages of the present application will be more apparent and understandable. The present application will be described in further detail below with reference to the drawings and specific embodiments.
[0018] In an exemplary embodiment, as shown in Figure 1 and Figure 2 , a cross-departmental equipment naming consistency evaluation method is provided, which comprises: S101, according to the multi-label annotation data of the equipment samples, a semantic label matrix is constructed by using a label splicing vectorization method.
[0019] According to the multi-label annotation data of the equipment samples, a dataset is constructed, which has equipment samples, categories, annotators. Each equipment sample can be annotated by multiple annotators, and the annotation of an equipment sample by an annotator is represented as a label set : .
[0020] A multi-annotator label splicing vector is constructed by using the formula .
[0021] wherein, is the label vector of the equipment sample annotated by the annotator , if , then , otherwise, the value is 0; represents the splicing operation of the vector; for example, the vector , the vector , and the splicing of the two vectors forms a new vector with a dimension of : .
[0022] After generating the multi-annotator label splicing vector for all equipment samples, a semantic label matrix is constructed by using ; wherein, the superscript T is the transpose.
[0023] The semantic label matrix The semantic label matrix is input as the label semantic representation of the entire data set for subsequent hash mapping or consistency calculation.
[0024] In order to ensure that the generated semantic label matrix can be normally input into the subsequent hash mapping or consistency calculation process, in the generated semantic label matrix, any illegal label will be detected and a skip strategy will be executed. For any illegal label (such as a label value or a non-integer type), a skip strategy is executed and zero is filled in the vector. For a missing value (such as a label not provided by a labeler), the label vector thereof is defined as: to maintain the dimension consistency and alignment of the spliced vector.
[0025] In S102, according to the semantic label matrix, the average similarity between different device samples is determined by using cosine similarity and Jaccard similarity; the similarity between the original samples is comprehensively evaluated by using the two indexes of cosine similarity and Jaccard similarity.
[0026] The cosine similarity between the device sample and the device sample is determined by using the formula .
[0027] wherein, is the dot product of two vectors, and are the Euclidean norms of and respectively. The range of the similarity measurement value is , 1 indicates complete similarity, 0 indicates no similarity, and -1 indicates complete opposition.
[0028] The Jaccard similarity between the device sample and the device sample is determined by using the formula .
[0029] wherein, and represent the semantic label matrix of the device sample and the device sample respectively.
[0030] The Jaccard similarity and the cosine similarity are combined by weighted average, the weight size of the two indexes is controlled by the weighted coefficient , and the average similarity between the two samples is obtained as: .
[0031] S103, according to the semantic label matrix, using standard local sensitive hashing algorithm and privacy protection mechanism, determine the hash signature between different device samples; the privacy protection mechanism is to flip the hash signature bit randomly according to the privacy level of the device.
[0032] The standard LSH method is used to calculate the hash signature of each device sample, which maps high-dimensional data to low-dimensional space through random projection, and then calculates the hash signature of each device sample. Given an input vector (multi-annotator label splicing vector) , wherein, represents the label semantic vector of the device sample, which contains various attribute information of the device (device type, model, specification parameter, functional characteristic). The calculation process of standard hash is as follows: (1) Random projection generation: a hyperplane in high-dimensional space is randomly generated, which divides the data points into two parts, one side and the other side of the hyperplane. Through this hyperplane, the data is projected into one-dimensional space. For a given input vector , a random vector is selected, and the dot product of the input vector is calculated, and the calculation formula is as follows: .
[0033] , wherein, represents the sign function, which maps the result of dot product to a binary value, returns or .
[0034] (2) Multiple projection generation: in order to improve the effect of hash, multiple random projections are carried out, that is, multiple random vectors are used to generate multiple hash functions. Each hash function calculates the dot product between the input vector and the corresponding random vector, and then generates a hash bit through the sign function: .
[0035] (3) Generate hash signature: the hash value of each input vector is composed of multiple hash bits, and the hash value of each input vector is a dimensional binary vector hash signature , which represents the mapping of input vector in low-dimensional space: .
[0036] In order to effectively prevent the leakage of original sensitive equipment data in the process of cross-department sharing, according to the privacy protection level a random vector is generated to flip bits in the hash signature . .
[0037] where denotes the bitwise XOR operation, is a random vector indicating the bits to be flipped. Each element of is generated by an independent Bernoulli random variable with probability .
[0038] S104, determine the similarity of hash signatures between different device samples, and correct the similarity using the privacy level. That is, compare the average similarity with the similarity after combining the hash signatures, and analyze whether the evaluation of the similarity after adding privacy protection has an impact; for the similarity error introduced by privacy protection, effectively restore the true similarity level on the basis of ensuring data irreversibility, and improve the practicality and credibility of the results.
[0039] S104 specifically includes: determine the similarity of hash signatures between different device samples using the formula . .
[0040] If two identical samples have probability of different bits, and different samples have probability of different bits, then the similarity of the corrected hash signature is determined using the formula .
[0041] where and denote the privacy-protected hash values of two device samples or annotators, is the number of bits of the hash signature, is an indicator function, which is 1 if , and 0 otherwise, used to count the matching of two hash signatures at each hash bit, is the privacy protection level, indicating how many hash bits are randomly flipped.
[0042] S104 further includes: performing consistency comparison to determine whether different devices remain consistent in semantic features; then performing clustering or grouping operations based on the corrected similarity matrix to realize the organization and feature aggregation of device samples; finally, these results can be used as input for subsequent anomaly detection and behavior modeling, thereby improving the practicality and credibility of the method.
[0043] Based on the same inventive concept, the embodiments of the present application also provide a cross-department device naming consistency evaluation system for implementing the above-mentioned cross-department device naming consistency evaluation method. The implementation scheme for solving the problem provided by the system is similar to the implementation scheme described in the above-mentioned method, and therefore the specific limitations in one or more cross-department device naming consistency evaluation system embodiments provided below can be referred to the limitations of the cross-department device naming consistency evaluation method described above, which will not be described here again.
[0044] In an exemplary embodiment, a cross-department device naming consistency evaluation system is provided, comprising: a label semantic vectorization module configured to construct a semantic label matrix by using a label concatenation vectorization method according to multi-label annotation data of device samples; a consistency evaluation module configured to determine an average similarity between different device samples by using cosine similarity and Jaccard similarity according to the semantic label matrix; a privacy protection hash matching module configured to determine hash signatures between different device samples by using a standard local sensitive hash algorithm and a privacy protection mechanism according to the semantic label matrix; the privacy protection mechanism is to randomly flip hash signature bits according to the privacy level of the device; a consistency evaluation and privacy protection similarity correction module configured to determine the similarity of hash signatures between different device samples and correct the similarity by using the privacy level.
[0045] As shown in Figure 3 the label semantic vectorization module is configured to convert multi-label annotation data into high-dimensional semantic vectors to support subsequent consistency evaluation. The input is multi-label annotation data anno_data provided by multiple annotators, which is converted into a unified semantic label matrix, i.e., a high-dimensional semantic label matrix vectors, by data structure analysis and label concatenation vectorization method after error detection, for subsequent hash mapping or consistency calculation.
[0046] The function of the consistency evaluation module is to evaluate the consistency of the generated semantic label matrix vectors, and to calculate the average similarity between different device samples by using cosine similarity and Jaccard similarity. The input is the semantic label matrix vectors output by the label semantic vectorization module. The similarity average value between samples is obtained by weighted average of cosine similarity and Jaccard similarity.
[0047] The privacy protection hash matching module function is to perform local sensitive hash (LSH) processing on the label semantic matrix, and introduce a privacy protection mechanism to avoid leaking sensitive information by perturbing the hash signature. The input is the semantic label matrix vectors output by the label semantic vectorization module. After multiple random projections, the dot product between the data points and the random vectors is calculated to obtain the hash signature. The output is the hash signature privacy_hashed_signatures calculated using local sensitive hash (LSH). The LSH algorithm is introduced to map high-dimensional label data to a low-dimensional hash space, greatly reducing the time complexity of similarity calculation between samples, and having good calculation performance and real-time response ability.
[0048] The consistency evaluation and privacy protection similarity correction module function is to calculate the similarity between hash signatures with privacy protection, and correct the errors introduced by privacy protection. The input is the hash signature privacy_hashed_signatures generated by the privacy protection hash matching module, the similarity calculation formula is used, and the privacy level is The similarity calculation value is corrected, and the output is the similarity value corrected_similarity calculated by the privacy protection similarity correction formula and the error caused by privacy protection is corrected.
[0049] The consistency evaluation index combination in the present application can be flexibly extended according to actual needs, such as introducing Hamming distance, Manhattan distance or point mutual information as supplementary indexes to adapt to different types of label data characteristics. The local sensitive hash mechanism used can be replaced by SimHash, MinHash or other efficient hash algorithms in specific scenarios to improve the adaptability and performance of different types of similarity tasks. It can be extended to other data governance fields that require multi-label integration, semantic consistency judgment and privacy protection It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant regulations.
[0050] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0051] The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a blockchain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.
[0052] In the present application, all actions of obtaining signals, information or data are performed under the premise of complying with the corresponding data protection regulations and policies of the country where the device is located, and under the premise of obtaining authorization from the owner of the corresponding device.
[0053] The technical features of the above embodiments can be combined arbitrarily. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, but as long as the combinations of the technical features do not exist, they should be considered as the scope of the present application.
[0054] The principles and implementation manners of the present application are described herein by using specific examples, and the above examples are only used to help understand the method of the present application and its core idea; meanwhile, for those skilled in the art, according to the idea of the present application, the specific implementation manners and application ranges will have changes. In conclusion, the content of the specification should not be understood as a limitation of the present application.
Claims
1. A method for assessing cross-departmental equipment naming consistency, characterized in that, The cross-department equipment naming consistency evaluation method comprises: According to the multi-label annotation data of the equipment samples, a label splicing vectorization method is used to construct a semantic label matrix; According to the semantic label matrix, cosine similarity and Jaccard similarity are used to determine the average similarity between different equipment samples; According to the semantic label matrix, a standard local sensitive hashing algorithm and a privacy protection mechanism are used to determine the hash signature between different equipment samples; the privacy protection mechanism is to randomly flip the hash signature bit according to the privacy level of the equipment; The similarity of the hash signatures between different equipment samples is determined, and the similarity is corrected by using the privacy level.
2. The cross-departmental device naming consistency assessment method of claim 1, wherein, According to the multi-label annotation data of the equipment samples, a label splicing vectorization method is used to construct a semantic label matrix, which specifically comprises: Using the formula Constructing multi-annotator label stitching vectors ; Utilizing Constructing a semantic tag matrix ; wherein, is the number of annotators is the label vector of the device sample, is the number of device samples, is the number of annotators, is the class of the device sample, is the number of device samples, denotes the concatenation operation of vectors, and the superscript T denotes the transpose.
3. The cross-departmental device naming consistency assessment method of claim 1, wherein, According to the multi-label annotation data of the equipment samples, a label splicing vectorization method is used to construct a semantic label matrix, which specifically comprises: Error detection is performed on the multi-label annotation data.
4. The cross-departmental device naming consistency assessment method of claim 1, wherein, According to the semantic label matrix, cosine similarity and Jaccard similarity are used to determine the average similarity between different equipment samples, which specifically comprises: According to the semantic label matrix, the cosine similarity and Jaccard similarity between two equipment samples are determined respectively; According to the cosine similarity and Jaccard similarity between two equipment samples, a weighted average method is used to determine the average similarity.
5. The cross-departmental device naming consistency assessment method of claim 4, wherein, According to the cosine similarity and Jaccard similarity between two equipment samples, a weighted average method is used to determine the average similarity, which specifically comprises: The average similarity between device samples is determined using the formula ; wherein, is a weighting factor, is a device sample and a cosine similarity between device samples is a Jaccard similarity between device samples is a device sample and a Jaccard similarity between device samples is a device sample 6. The cross-departmental device naming consistency assessment method of claim 1, wherein, According to the semantic label matrix, a standard local sensitive hashing algorithm and a privacy protection mechanism are used to determine the hash signature between different equipment samples, which specifically comprises: Using the formula determining the hash signature of the device sample i ; According to a privacy protection level Generating a vector of hash bits in a random flip hash signature ; Using the formula Determining the flipped hash signature ; wherein, is the hash bit generated after k random projections for device sample i, denotes a bitwise XOR operation, and the random vector is generated by independent Bernoulli random variables with probability privacy level.
7. The cross-departmental device naming consistency assessment method of claim 1, wherein, The similarity of the hash signatures between different equipment samples is determined, and the similarity is corrected by using the privacy level, which specifically comprises: Using the formula Determining similarity of hash signatures between different device samples ; Using the formula Determining similarity of corrected hash signatures ; wherein, and denote the privacy-protected hash values of two device samples or annotators, is the number of bits of the hash signature, is an indicator function, which is 1 if and 0 otherwise, used to count the matching of two hash signatures at each hash bit, is the privacy protection level, indicating how many hash bits are flipped randomly.
8. A cross-departmental device naming consistency assessment system, comprising: The cross-department equipment naming consistency evaluation system comprises: A label semantic vectorization module is configured to construct a semantic label matrix according to multi-label annotation data of equipment samples by using a label splicing vectorization method; A consistency evaluation module is configured to determine the average similarity between different equipment samples according to the semantic label matrix by using cosine similarity and Jaccard similarity; A privacy protection hash matching module is configured to determine the hash signature between different equipment samples according to the semantic label matrix by using a standard local sensitive hashing algorithm and a privacy protection mechanism; the privacy protection mechanism is to randomly flip the hash signature bit according to the privacy level of the equipment; A consistency evaluation and privacy protection similarity correction module is configured to determine the similarity of the hash signatures between different equipment samples, and correct the similarity by using the privacy level.
Citation Information
Patent Citations
Personalized news recommendation device and method based on news content and theme feature
CN102831234A
An approximate member query method based on hamming distance
CN109034197A
MES-oriented mass data redundancy elimination method and system
CN112162977A
Depth cross-modal hash image retrieval method based on joint semantic matrix
CN113177132A
Deep cross-modal hashing method based on fusion similarity
CN114359930A