Semantic cognition-based iot terminal data classification method and application thereof

By constructing a data fusion computing model and semantic cognition technology, the problem of data silos in IoT terminal devices has been solved, and semantic annotation and unification of data have been achieved, thereby improving the intelligence level of the devices.

CN116933128BActive Publication Date: 2026-05-01JIANGSU ZHONGRUN PUDA ENVIRONMENTAL BIG DATA CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JIANGSU ZHONGRUN PUDA ENVIRONMENTAL BIG DATA CO LTD
Filing Date
2023-07-10
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

IoT sensing devices generate a large amount of isolated and heterogeneous sensing data, forming data silos, which leads to data heterogeneity and difficulties in upgrading devices to be intelligent.

Method used

A semantic cognition-based IoT terminal data classification method constructs a data fusion computing model, utilizes LSTM neural network model and LDA topic model for feature enhancement and data fusion, and employs Adaboost algorithm to optimize the classifier, thereby achieving semantic annotation and unification of the data.

Benefits of technology

It effectively solves the data silo problem of IoT terminal devices, realizes the unification of semantic information and the intelligent upgrade of devices, and improves the accuracy and efficiency of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116933128B_ABST
    Figure CN116933128B_ABST
Patent Text Reader

Abstract

The embodiment of the application discloses a kind of based on semantic cognition's Internet of Things terminal data classification method and its application, wherein based on semantic cognition's Internet of Things terminal data classification method includes: according to the semantic feature vector information of the reinforced application scene and business of Internet of Things terminal equipment, data fusion calculation model is constructed to carry out data fusion;Through the classification model trained based on the fusion feature of multi-level feature fusion, carry out cognitive calculation, classify the data fused using the data fusion calculation model, obtain the final application scene and the classification and result of associated business.It solves the problem that a large number of isolated and heterogeneous perception data are generated by Internet of Things perception device in prior art at all times, and a large number of data islands are formed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer information processing technology, specifically to a semantic cognition-based method for classifying IoT terminal data and its application. Background Technology

[0002] With the emergence of the Internet of Things (IoT) concept, IoT systems, which incorporate a large number of sensing devices, are increasingly being used in various fields. However, these sensing devices constantly generate a large amount of isolated and heterogeneous sensing data, creating numerous data silos. Summary of the Invention

[0003] The purpose of this application is to provide a semantic cognition-based IoT terminal data classification method and its application, in order to solve the problem that IoT sensing devices in the prior art constantly generate a large amount of isolated and heterogeneous sensing data, forming a large number of data silos.

[0004] To achieve the above objectives, this application provides a semantic cognition-based IoT terminal data classification method, including: constructing a data fusion calculation model to perform data fusion based on the enhanced semantic feature vector information of the IoT terminal device application scenario and business;

[0005] A classification model trained based on multi-level feature fusion is used to perform cognitive computing. The data fused by the data fusion computing model is then classified to obtain the final application scenarios, related business classifications, and results.

[0006] Optionally, the step of constructing a data fusion computing model based on the enhanced semantic feature vector information of the IoT terminal device application scenario and business includes:

[0007] Obtain sample data of IoT terminal device application scenarios;

[0008] Based on the application scenarios and related business characteristics of IoT terminal devices, the sample data is preprocessed to extract corresponding semantic feature vectors and establish rules for the application scenario and related business ontology library.

[0009] The semantic feature vector is enhanced by a neural network model to obtain the enhanced semantic feature vector.

[0010] Optionally, the preprocessing includes ontology parsing and text processing; wherein

[0011] The ontology parsing includes parsing the concept information in the initial ontology library to obtain a concept set; parsing the instance information in the initial ontology library to obtain an instance set; and parsing the relation information in the initial ontology library to obtain a relation set.

[0012] The text processing includes webpage text processing and collected data text processing; the webpage text processing includes: denoising the webpage and extracting webpage features; the collected data text processing includes: converting the data format.

[0013] Optionally, the step of strengthening the semantic feature vector through a neural network model to obtain the strengthened semantic feature vector includes: inputting the semantic feature vector into an LSTM neural network model for feature strengthening training; wherein

[0014] The LSTM neural network model includes an input gate, a forget gate, and an output gate;

[0015] The input gate saves all the information input at the current time into the cell state at the current time, and calculates the candidate information at the current time through the tanh function, and calculates the decision vector at the current time, thereby determining the amount of information input into the cell state at the current time.

[0016] The forget gate saves the cell state from the previous time step to the cell state at the current time step;

[0017] The output gate controls the amount of information that is ultimately output as the current state of the unit.

[0018] Optionally, the step of constructing a data fusion computing model to perform data fusion includes:

[0019] LDA topic modeling is used to perform modeling analysis and identify specific business in specific scenarios, identify new business in specific scenarios, compare detection indicators with thresholds, and then derive new business in specific scenarios.

[0020] Optionally, the construction of the data fusion computing model for data fusion specifically includes:

[0021] Data modeling is performed using Gibbs sampling to solve for the latent variables θ, φ, z, thereby obtaining the joint probability distribution p(θ, φ, z | w, α, β);

[0022] For model training, based on the number of keywords provided by experts, the hyperparameter of the LDA topic model is set to the number of topics K, and the hyperparameter search range of the number of topics K is set.

[0023] For model evaluation, the coherence score was selected as the model evaluation index to evaluate and verify the K value.

[0024] In K given topics, when at least one pair of key information in any two topics is more correlated than a preset correlation threshold, the two topics are merged. The topic with the lower coherence value is merged into the topic with the higher coherence value, resulting in multiple emerging topics of number K', which yields new businesses in specific scenarios.

[0025] Optionally, the classification model is based on the AdaBoost algorithm, uses a BP neural network as a weak classifier for the AdaBoost algorithm, and employs a symbiotic biological search algorithm to optimize the weights of each weak classifier.

[0026] Optionally, the training of the classification model includes training of key features and training of fused features; for each type of key feature corresponding to a scenario, a weak classifier is trained to achieve text classification and numerical classification; for each type of fused feature corresponding to a specific business, multiple weak classifiers are trained to achieve multimodal fusion classification; and all weak classifiers are combined to obtain the AdaBoost ensemble classifier.

[0027] To achieve the above objectives, this application also provides a semantic cognition-based IoT terminal data classification device, comprising: a memory; and

[0028] A processor connected to the memory, the processor being configured to perform the steps of the method described above.

[0029] To achieve the above objectives, this application also provides a computer storage medium having a computer program stored thereon, wherein the computer program, when executed by a machine, implements the steps of the method described above.

[0030] The embodiments of this application have the following advantages:

[0031] This application provides a semantic cognition-based IoT terminal data classification method, including: constructing a data fusion computing model to perform data fusion based on the enhanced semantic feature vector information of the IoT terminal device application scenario and business; performing cognitive computing through a classification model trained on the fusion features after multi-level feature fusion, classifying the data fused using the data fusion computing model, and obtaining the final application scenario and related business classification and results.

[0032] By using the above methods, different devices and their generated data information are semantically labeled with scenarios and related businesses, thereby constructing data association models in different domains. This helps to shield data heterogeneity, achieve semantic information unification, and effectively solve the problems of a large number of data silos formed by IoT terminal devices and the intelligent upgrading and transformation of devices. Attached Figure Description

[0033] To more clearly illustrate the embodiments of this application or the technical solutions in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely exemplary, and those skilled in the art can derive other embodiments based on the provided drawings without creative effort.

[0034] Figure 1 A flowchart illustrating a semantic cognition-based IoT terminal data classification method provided in this application embodiment;

[0035] Figure 2 This is a block diagram of a semantic cognition-based IoT terminal data classification device provided in an embodiment of this application. Detailed Implementation

[0036] The following specific embodiments illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0037] Furthermore, the technical features involved in the different embodiments of this application described below can be combined with each other as long as they do not conflict with each other.

[0038] One embodiment of this application provides a method for classifying IoT terminal data based on semantic cognition, referencing... Figure 1 , Figure 1 The flowchart provided in one embodiment of this application illustrates a semantic cognition-based IoT terminal data classification method. It should be understood that the method may also include additional boxes not shown and / or the boxes shown may be omitted, and the scope of this application is not limited in this respect.

[0039] In step 101, a data fusion computing model is constructed to perform data fusion based on the enhanced semantic feature vector information of the IoT terminal device application scenario and business.

[0040] Specifically, the model building module constructs a data fusion computing model based on the enhanced semantic feature vector information of IoT terminal device application scenarios and services to perform data fusion, thereby masking data heterogeneity and achieving unified semantic information and intelligent application. Here, the LDA topic model is mainly used for modeling analysis and identification of topics (specific services in specific scenarios), identification of emerging topics (new services in specific scenarios), comparison of detection indicators with thresholds, and finally, the emergence of emerging topics (new services in specific scenarios).

[0041] In some embodiments, constructing a data fusion computing model based on the enhanced semantic feature vector information of the IoT terminal device application scenario and services includes:

[0042] 1. Obtain sample data of IoT terminal device application scenarios through the data acquisition unit;

[0043] 2. Through the semantic extraction unit, based on the application scenario and related business characteristic information of the IoT terminal device, the sample data is preprocessed to extract the corresponding semantic feature vectors, and rules for the application scenario and related business ontology library are established. In some embodiments,

[0044] The preprocessing includes ontology parsing and text processing.

[0045] The ontology parsing includes parsing the concept information in the initial ontology library to obtain a concept set; parsing the instance information in the initial ontology library to obtain an instance set; and parsing the relation information in the initial ontology library to obtain a relation set.

[0046] The text processing includes processing webpage text and processing collected data text. Webpage text processing includes: denoising the webpage and extracting its features; the collected data text processing includes: converting the data format.

[0047] It also includes updating and expanding the initial ontology library. Updating the initial ontology library includes the following processes: when real-time dynamic measurement values ​​sent by the sensor are collected, dynamic data matching is performed to update the instance set in the initial ontology library; when inherent attribute information values ​​sent by the sensor are collected, static data matching is performed to update the concept set and relation set in the initial ontology library.

[0048] The expansion of the initial ontology database includes the following processes: filtering information collected from the network to obtain network information with high similarity to the ontology database; and obtaining highly relevant vocabulary information by calculating edit distance and context similarity. If the relevance of a document to the sensor ontology of an IoT terminal device is greater than the domain document relevance coefficient of 0, then the domain document relevance is high; otherwise, the domain document relevance is low.

[0049] The calculation of edit distance and context similarity includes: given two words A and B, the concept matching degree of the two words is obtained according to the Sigmoid function, as well as the edit distance and context similarity. Here, the Sigmoid function is used. If the concept matching degree Sim_rept(t1, t2) of the two words is greater than the domain word relevance coefficient, then the domain word relevance is high; otherwise, the domain word relevance is low.

[0050] 3. The semantic feature vector is enhanced using a neural network model through a semantic enhancement unit, resulting in an enhanced semantic feature vector. Specifically, the feature vector enhancement unit is used to establish an LSTM neural network model, inputting the semantic feature vector into the LSTM neural network model for feature enhancement training. In some embodiments,

[0051] The LSTM neural network model includes an input gate, a forget gate, and an output gate;

[0052] The input gate saves all the information input at the current time into the cell state at the current time, and calculates the candidate information at the current time through the tanh function, and calculates the decision vector at the current time, thereby determining the amount of information input into the cell state at the current time.

[0053] The forget gate saves the cell state from the previous time step into the cell state at the current time step; the output gate controls the amount of information that is finally output from the cell state at the current time step.

[0054] In some embodiments, the LSTM neural network model further includes a hidden layer; the input gate saves all the information input at the current time into the cell state at the current time, calculates the decision vector at the current time, and calculates the candidate information at the current time through the tanh function, thereby determining the amount of information input into the cell state at the current time, specifically including:

[0055] At time t, the input gate controls the input value x at the current time. t Calculate the decision vector i of the input gate at the current time. t Its calculation formula is

[0056] i t =σ(w ix x t +w ih h t-1 +b i )

[0057] Among them, h t-1 Let w be the hidden state of the hidden layer at time t-1. ix After passing through the weight parameter matrix of the input gate, w ih Let σ be the weight parameter matrix between the input gate and the hidden layer, and b be the weight parameter matrix between the input gate and the hidden layer. i This is the bias term for the input gate;

[0058] Calculate the candidate information c% at the current time using the tanh function. t Its calculation formula is

[0059] c% t=tanh(w cx x t +w ch h t-1 +b c )

[0060] Among them, w cx Let w be the weight parameter matrix after the output gate. ch Let be the weight parameter matrix between the output gate and the hidden layer, tanh be the hyperbolic tangent function, and b be the weight parameter matrix between the output gate and the hidden layer. c This is the bias term for the output gate;

[0061] c% of the candidate information t The decision vector i of the input gate t Multiplication determines the amount of information stored in the state unit.

[0062] In some embodiments, the forget gate saving the cell state from the previous time step to the cell state at the current time step specifically includes:

[0063] Determine the cell state c at the previous moment. t-1 The cell state c added to the current time step t The amount of information in the memory is used to calculate the activation vector value f of the forget gate at time t. t Its calculation formula is

[0064] f t =σ(w fx x t +w fh h t-1 +b f )

[0065] Among them, w fx Let w be the weight parameter matrix after passing through the forget gate. fh Let b be the weight parameter matrix between the forget gate and the hidden layer. f For the bias term of the forget gate;

[0066] Using activation vector value f t Multiply by the cell state at the previous time step c t-1 Determine the cell state c from the previous moment. t-1 The cell state c added to the current time step t The amount of information contained within.

[0067] In some embodiments, the amount of information whose current cell state is ultimately output by the output gate specifically includes:

[0068] Calculate the decision vector o of the output gate at the current time. t Its calculation formula is

[0069] ot =σ(w ox x t +w oh h t-1 +b o )

[0070] Among them, w ox Let w be the weight parameter matrix after the output gate. oh Let b be the weight parameter matrix between the output gate and the hidden layer. o This is the bias term for the output gate;

[0071] Calculate the hidden state h of the hidden layer at the current time step. t Its calculation formula is

[0072] h t =o t ·tanh(c t )

[0073] Among them, c t This represents the current state of the cell.

[0074] The decision vector o of the output gate t Multiply by the hidden state at the current time by h t This determines the amount of information that needs to be output at the current moment.

[0075] In some embodiments, based on the above-described data fusion computing model construction scheme, the data fusion process specifically includes:

[0076] S4-1, Data Modeling: Using Gibbs Sampling to Solve for Latent Variable θ z, and thus the joint probability distribution is obtained. The LDA topic model uses the third-party LdaMallet algorithm, specifying the bag-of-words model of the text corpus, the number of topics, the hyperparameter α, id2word, the number of concurrent processor cores (workers), and the number of training iterations; then proceed to step S4-2.

[0077] S4-2, Model Training: Based on the number of keywords provided by experts, set the hyperparameter of the LDA topic model to the number of topics K, and set the hyperparameter search range for the number of topics K; continue to step S4-3.

[0078] S4-3, Model Evaluation: The coherence score is used as the model evaluation metric to evaluate and validate the K value. Coherence is a measure of the interpretability of a topic to humans. Given keywords provided by experts as a given topic, and considering the word distribution matrix of that topic, the coherence value for any given topic is:

[0079]

[0080] Where w a ,w b All words are from the word distribution matrix, i.e., related words for a given topic. The top n most frequent related words for each given topic in the corpus, ranked from highest to lowest frequency, are used as key information. D(w) i ) is defined as containing the word w i The number of documents, D(w) i ,w j ) is defined as simultaneously containing the word w i and w j The number of documents, D, is defined as the total number of document information in the corpus. The key information w within any two topics is calculated. i and w j Interrelationships:

[0081]

[0082] Continue with step S4-4;

[0083] S4-4; In K given topics, when at least one pair of key information in any two topics is more correlated than a preset correlation threshold, the two topics are merged, and the topic with the lower coherence value is merged into the topic with the higher coherence value, resulting in multiple emerging topics of number K′.

[0084] In step 102, a cognitive computation is performed using a classification model trained on fused features after multi-level feature fusion. The data fused using the data fusion computation model is then classified to obtain the final application scenario, related business classification, and results.

[0085] Specifically, through the cognitive computing unit, a classification model is trained based on the fused features after multi-level feature fusion. Cognitive computing is then performed through the classification model to classify the data fused using the data fusion computing model, thereby obtaining the final application scenario and related business classification and results.

[0086] In some embodiments, the classification model is based on the AdaBoost algorithm, uses a BP neural network as a weak classifier for the AdaBoost algorithm, and employs a symbiotic biological search algorithm to optimize the weights of each weak classifier.

[0087] In some embodiments, the training of the classification model includes training of key features and training of fused features; for each type of key feature corresponding to a scenario, a weak classifier is trained to achieve text classification and numerical classification; for each type of fused feature corresponding to a specific business, multiple weak classifiers are trained to achieve multimodal fusion classification; and all weak classifiers are combined to obtain the AdaBoost ensemble classifier.

[0088] In some embodiments, the Adaboost algorithm can be divided into three steps: 1) First, the weight distribution D1 of the training data is initialized. Assuming there are N training samples, each training sample is initially assigned the same weight: w i =1 / N; 2) Then, train the weak classifier hi. The specific training process is as follows: if a training sample point is accurately classified by the weak classifier hi, then its corresponding weight should be reduced in the next training set; conversely, if a training sample point is misclassified, then its weight should be increased. The sample set with updated weights is used to train the next classifier, and the entire training process continues iteratively; 3) Finally, combine the various trained weak classifiers into a strong classifier. After the training process of each weak classifier is completed, the symbiotic biological search algorithm is used to optimize the weights of each weak classifier, further increasing the weight of the weak classifier with a small classification error rate, so that it plays a larger decisive role in the final classification function, while reducing the weight of the weak classifier with a large classification error rate, so that it plays a smaller decisive role in the final classification function. In other words, the weak classifier with a low error rate has a larger weight in the final classifier, and so on.

[0089] Then, based on the adjusted sample distribution, the next base learner is trained. This process is repeated until the number of base learners reaches a pre-specified value T. Finally, these T learners are weighted and combined. The training set remains the same in each round; only the weight of each example in the training set changes in the classifier. The weights are adjusted based on the classification results of the previous round, continuously adjusting the weights of the examples according to the error rate. The higher the error rate, the greater the weight. Each weak classifier has a corresponding weight, with classifiers with smaller classification errors receiving larger weights. The prediction functions can only be generated sequentially because the parameters of the later model require the results of the previous model.

[0090] The above methods employ semantic cognition to describe the data resources generated by IoT terminal devices, resulting in highly readable descriptions that are easily understood and processed by the IoT platform. Simultaneously, resource requests from IoT applications are also based on semantic descriptions, enabling the IoT platform to accurately perform corresponding cognitive calculations and quickly process problems based on the resources of IoT applications. By semantically labeling different devices and their generated data with scenarios and related services, data association models across different domains are constructed to mask data heterogeneity, achieve semantic information unification, and effectively solve the problems of numerous data silos formed by IoT terminal devices and the challenges of intelligent upgrades and transformations of these devices.

[0091] Figure 2 A block diagram of a semantic cognition-based IoT terminal data classification device provided in this application embodiment. The device includes:

[0092] The memory 201; and the processor 202 connected to the memory 201, the processor 202 being configured to: construct a data fusion computing model to perform data fusion based on the enhanced semantic feature vector information of the IoT terminal device application scenario and business;

[0093] A classification model trained based on multi-level feature fusion is used to perform cognitive computing. The data fused by the data fusion computing model is then classified to obtain the final application scenarios, related business classifications, and results.

[0094] In some embodiments, the processor 202 is further configured to: construct a data fusion computing model based on the enhanced semantic feature vector information of the IoT terminal device application scenario and services, including:

[0095] Obtain sample data of IoT terminal device application scenarios;

[0096] Based on the application scenarios and related business characteristics of IoT terminal devices, the sample data is preprocessed to extract corresponding semantic feature vectors and establish rules for the application scenario and related business ontology library.

[0097] The semantic feature vector is enhanced by a neural network model to obtain the enhanced semantic feature vector.

[0098] In some embodiments, the processor 202 is further configured such that the preprocessing includes ontology parsing and text processing; wherein

[0099] The ontology parsing includes parsing the concept information in the initial ontology library to obtain a concept set; parsing the instance information in the initial ontology library to obtain an instance set; and parsing the relation information in the initial ontology library to obtain a relation set.

[0100] The text processing includes webpage text processing and collected data text processing; the webpage text processing includes: denoising the webpage and extracting webpage features; the collected data text processing includes: converting the data format.

[0101] In some embodiments, the processor 202 is further configured to: enhance the semantic feature vector through a neural network model to obtain the enhanced semantic feature vector, comprising: inputting the semantic feature vector into an LSTM neural network model for feature enhancement training; wherein

[0102] The LSTM neural network model includes an input gate, a forget gate, and an output gate;

[0103] The input gate saves all the information input at the current time into the cell state at the current time, and calculates the candidate information at the current time through the tanh function, and calculates the decision vector at the current time, thereby determining the amount of information input into the cell state at the current time.

[0104] The forget gate saves the cell state from the previous time step to the cell state at the current time step;

[0105] The output gate controls the amount of information that is ultimately output as the current state of the unit.

[0106] In some embodiments, the processor 202 is further configured to: construct the data fusion computing model to perform data fusion, including:

[0107] LDA topic modeling is used to perform modeling analysis and identify specific business in specific scenarios, identify new business in specific scenarios, compare detection indicators with thresholds, and then derive new business in specific scenarios.

[0108] In some embodiments, the processor 202 is further configured to: construct the data fusion computing model to perform data fusion, specifically including:

[0109] Data modeling was performed using Gibbs sampling to solve for the latent variable θ. z, and thus the joint probability distribution is obtained.

[0110] For model training, based on the number of keywords provided by experts, the hyperparameter of the LDA topic model is set to the number of topics K, and the hyperparameter search range of the number of topics K is set.

[0111] For model evaluation, the coherence score was selected as the model evaluation index to evaluate and verify the K value.

[0112] In K given topics, when at least one pair of key information in any two topics is more correlated than a preset correlation threshold, the two topics are merged. The topic with the lower coherence value is merged into the topic with the higher coherence value, resulting in multiple emerging topics of number K', which yields new businesses in specific scenarios.

[0113] In some embodiments, the processor 202 is further configured such that: the classification model is based on the AdaBoost algorithm, a BP neural network is used as a weak classifier for the AdaBoost algorithm, and a symbiotic biological search algorithm is used to optimize the weights of each weak classifier.

[0114] In some embodiments, the processor 202 is further configured to: train the classification model including key feature training and fusion feature training; train a weak classifier for each key feature type corresponding to each scenario to achieve text classification and numerical classification; train multiple weak classifiers for each fusion feature corresponding to each specific business to achieve multimodal fusion classification; and combine all weak classifiers to obtain an AdaBoost ensemble classifier.

[0115] For specific implementation methods, please refer to the aforementioned method embodiments, which will not be repeated here.

[0116] This application may be a method, apparatus, system, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of this application.

[0117] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0118] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0119] The computer program instructions used to perform the operations of this application may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuits, such as programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), are personalized by utilizing state information from the computer-readable program instructions. These electronic circuits can execute the computer-readable program instructions to implement various aspects of this application.

[0120] Various aspects of this application are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0121] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0122] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0123] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0124] Note that, unless otherwise explicitly stated, all features disclosed in this specification (including any appended claims, abstract, and drawings) may be replaced by alternative features for achieving the same, equivalent, or similar purpose. Therefore, unless explicitly stated otherwise, each disclosed feature is merely one example of a set of equivalent or similar features. Where used, "further," "preferably," "even further," and "more preferably" are simple starting points for describing another embodiment based on the foregoing embodiments, the combination of which with the foregoing embodiments constitutes the complete configuration of another embodiment. Any combination of several "further," "preferably," "even further," or "more preferably" settings following the same embodiment constitutes yet another embodiment.

[0125] Although this application has been described in detail above with general descriptions and specific embodiments, some modifications or improvements can be made to it, which will be obvious to those skilled in the art. Therefore, all such modifications or improvements made without departing from the spirit of this application fall within the scope of protection claimed in this application.

Claims

1. A method for classifying IoT terminal data based on semantic cognition, characterized in that, include: Based on the enhanced semantic feature vector information of IoT terminal device application scenarios and services, a data fusion computing model is constructed to perform data fusion. Cognitive computing is performed by using a classification model trained based on multi-level feature fusion to classify the data fused by the data fusion computing model, thereby obtaining the final application scenarios, related business classifications, and results. The construction of the data fusion computing model for data fusion specifically includes: Data modeling is performed using Gibbs sampling to solve for the latent variables θ, φ, z, thereby obtaining the joint probability distribution p(θ, φ, z | w, α, β); For model training, based on the number of keywords provided by experts, the hyperparameter of the LDA topic model is set to the number of topics K, and the hyperparameter search range of the number of topics K is set. For model evaluation, the coherence score was selected as the model evaluation index to evaluate and verify the K value. Among K given topics, when any two topics have at least one pair of key information points with a correlation higher than a preset correlation threshold, these two topics are merged. Topics with lower coherence values ​​are merged into topics with higher coherence values, resulting in a quantity of K. Several emerging themes have led to new businesses in specific scenarios.

2. The IoT terminal data classification method based on semantic cognition according to claim 1, characterized in that, The step of constructing a data fusion computing model based on the enhanced semantic feature vector information of IoT terminal device application scenarios and services includes: Obtain sample data of IoT terminal device application scenarios; Based on the application scenarios and related business characteristics of IoT terminal devices, the sample data is preprocessed to extract corresponding semantic feature vectors and establish rules for the application scenario and related business ontology library. The semantic feature vector is enhanced by a neural network model to obtain the enhanced semantic feature vector.

3. The IoT terminal data classification method based on semantic cognition according to claim 2, characterized in that, The preprocessing includes ontology parsing and text processing; in The ontology parsing includes parsing the concept information in the initial ontology library to obtain a concept set; parsing the instance information in the initial ontology library to obtain an instance set; and parsing the relation information in the initial ontology library to obtain a relation set. The text processing includes text processing of web pages and text processing of collected data; The webpage text processing includes: removing noise from the webpage and extracting webpage features; the collected data text processing includes: converting the data format.

4. The IoT terminal data classification method based on semantic cognition according to claim 2, characterized in that, The step of enhancing the semantic feature vector using a neural network model to obtain the enhanced semantic feature vector includes: inputting the semantic feature vector into an LSTM neural network model for feature enhancement training; wherein... The LSTM neural network model includes an input gate, a forget gate, and an output gate; The input gate saves all the information input at the current time into the cell state at the current time, and calculates the candidate information at the current time through the tanh function, and calculates the decision vector at the current time, thereby determining the amount of information input into the cell state at the current time. The forget gate saves the cell state from the previous time step to the cell state at the current time step; The output gate controls the amount of information that is ultimately output as the current state of the unit.

5. The IoT terminal data classification method based on semantic cognition according to claim 4, characterized in that, The construction of the data fusion computing model for data fusion includes: LDA topic modeling is used to perform modeling analysis and identify specific business in specific scenarios, identify new business in specific scenarios, compare detection indicators with thresholds, and then derive new business in specific scenarios.

6. The IoT terminal data classification method based on semantic cognition according to claim 1, characterized in that, The classification model is based on the AdaBoost algorithm, uses a BP neural network as a weak classifier for the AdaBoost algorithm, and employs a symbiotic biological search algorithm to optimize the weights of each weak classifier.

7. The IoT terminal data classification method based on semantic cognition according to claim 6, characterized in that, The training of the classification model includes training of key features and training of fused features; For each type of key feature, a weak classifier is trained to achieve text classification and numerical classification; for each type of specific business, multiple weak classifiers are trained to achieve multimodal fusion classification; and all weak classifiers are combined to obtain the AdaBoost ensemble classifier.

8. A data classification device for Internet of Things (IoT) terminals based on semantic cognition, characterized in that, include: Memory; as well as A processor connected to the memory, the processor being configured to perform the steps of the method as claimed in any one of claims 1 to 7.

9. A computer storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a machine, it implements the steps of the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Multi-source heterogeneous data classification system and method based on semantic enhancement and feature fusion

    CN114170458A

  • Device for generating text classification model, method, and computer readable storage medium

    WO2019200806A1