Tin-indium material public opinion data management method and device, equipment and storage medium
By building a TiDB cluster and using MySQL client to manage the database, combining data crawler and emotional polarity classification technology, the unified problem of public opinion data management of tin indium materials is solved, and efficient data management and query are achieved.
Patent Information
- Application Number
- CN202510055463.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, it is difficult to uniformly manage public opinion data of tin indium materials from various channels, resulting in resistance to data integration, monitoring and management.
Build TiDB's global environment variable file through TiUP component manager, generate TiDB cluster, and connect to the cluster through MySQL client. Use data crawling scripts to crawl public web page data, perform feature extraction and emotional polarity classification, form a tin indium material public opinion database, and query, update and delete the database through SQL statements.
It realizes unified management of public opinion data of tin indium material, improves data queryability and management efficiency, and simplifies the database construction and operation process.
Smart Images

Figure CN119938644A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of electrical digital data processing, and in particular to a data management method, device, equipment and storage medium for public opinion on tin and indium materials. Background Art
[0002] With the rapid development of the Internet and information technology, especially new media such as online platforms and online media, their rich forms and high coverage have had a huge impact on political, economic, cultural and social life. The resulting online public opinion problems have become an important factor affecting social stability, and it is necessary to monitor, analyze and warn of public opinion.
[0003] As a widely used alloy metal in modern times, tin-indium material has an irreplaceable position in many industries. For example, adding indium to solder can significantly improve welding quality. Tin-indium alloy can reduce welding defects and porosity problems, lower welding temperature, and enhance welding strength and hardness. Tin-indium oxide is used to manufacture transparent conductive films in liquid crystal displays and mobile phone screens. Tin-indium alloy is used in the semiconductor field to manufacture high-efficiency polycrystalline silicon solar cells, which can improve the conversion efficiency of solar cells and reduce their cost. Tin-indium alloy can be used to manufacture high-temperature alloys and other metal alloys, further expanding its application range in industry. Tin-indium alloy can also be used to manufacture electromagnetic compatibility materials and various sensors to enhance the electromagnetic compatibility and sensing performance of equipment. Tin-indium alloy can be used to manufacture key components in medical devices to enhance the performance and reliability of medical devices.
[0004] Since tin indium materials have a wide range of uses and involve many industries, it is necessary to monitor the industry technology, development status, social discussions, etc. of tin indium materials in order to accurately judge the application prospects, technical shortcomings, and social evaluation of tin indium materials.
[0005] At present, the monitoring of industry technology, development status, social discussions, etc. based on tin and indium materials is still blank. Subjective and objective information such as the development trend, industry status, and industry evaluation of the tin and indium materials industry need to be manually searched, captured, monitored, and saved in real time. Due to differences in acquisition channels, the data formats obtained are different, which leads to varying degrees of resistance to the subsequent integration, monitoring, and management of various types and formats of data. Summary of the invention
[0006] The main purpose of this application is to provide a data management method, device, equipment and storage medium for tin indium material public opinion, so as to solve the problem in the prior art that it is difficult to uniformly manage tin indium material public opinion from various channels.
[0007] In order to achieve the above objectives, this application provides the following technical solutions:
[0008] A data management method for public opinion on tin-indium materials, wherein the public opinion on tin-indium materials appears on a public webpage of an alloy material industry, and the data management method comprises:
[0009] Step S1, build the global environment variable file of TiDB through the TiUP component manager;
[0010] Step S2, generating a TiDB cluster in the global environment variable file through the playground instruction of the TiUP component manager;
[0011] Step S3, connecting to the TiDB cluster via a MySQL client signal;
[0012] Step S4, crawling a number of public data related to tin and indium materials from the public webpage through a data crawler script;
[0013] Step S5, extracting features from all public data to obtain several keywords;
[0014] Step S6, classifying all keywords according to the sentiment polarity and occurrence timestamp, and obtaining a sentiment polarity classification table based on a preset time period;
[0015] Step S7, creating a blank database with the same dimension as the sentiment polarity classification table in the TiDB cluster through the SQL creation statement of the MySQL client;
[0016] Step S8, importing all sentiment polarity classification tables into the blank database through the SQL import statement of the MySQL client to form a tin and indium material public opinion database;
[0017] Step S9, query, update and delete data on the tin-indium material public opinion database through the SQL query statement, SQL update statement and SQL delete statement of the MySQL client.
[0018] As a further improvement of the present application, step S8, importing all sentiment polarity classification tables into the blank database through the SQL import statement of the MySQL client to form a tin-indium material public opinion database, then comprising:
[0019] Step S10, splitting the tin-indium material public opinion database into a plurality of sub-databases based on the preset time period;
[0020] Step S20, obtaining the chain hash certificate of each sub-database through the SHA256 algorithm;
[0021] Step S30, copy all the hash certificates on the chain and save them locally, so as to obtain a local hash certificate based on a sub-database;
[0022] Step S40, saving all the hash evidence on the blockchain in order from old to new according to the preset time period;
[0023] Step S50, based on the current local hash evidence, retrieve the on-chain hash evidence that matches the local hash evidence in the blockchain;
[0024] Step S60, determining whether the on-chain hash evidence is retrieved successfully, if the on-chain hash evidence is not retrieved successfully, executing step S70;
[0025] Step S70, determining whether one of the on-chain hash evidence and the local hash evidence has been tampered with;
[0026] Step S80: Generate a data tampering signal and send it to an external monitoring terminal.
[0027] As a further improvement of the present application, in step S5, all public data are subjected to feature extraction to obtain several keywords, including:
[0028] Step S51, performing word segmentation processing on each public data by using a word segmentation algorithm, and obtaining a data set to be processed having a plurality of word segmentation sub-data based on one public data;
[0029] Step S52: extract keywords of all word segmentation sub-data in the current data set to be processed by using a feature extraction algorithm.
[0030] As a further improvement of the present application, in step S51, each public data is segmented by a segmentation algorithm, and a data set to be processed having a plurality of segmentation sub-data is obtained based on the public data, including:
[0031] Step S511, performing word segmentation processing on each public data by using SentencePiece to mark the word segmentation point of each public data;
[0032] Step S512, dividing the current public data into a number of segmentations based on all segmentation points;
[0033] Step S513, marking each word segmentation by decoding the time tag, and obtaining a word segmentation sub-data based on one mark;
[0034] Step S514: Pack all the word segmentation sub-data of the current public data into a data set to be processed.
[0035] As a further improvement of the present application, step S52, extracting keywords of all word segmentation sub-data in the current data set to be processed by a feature extraction algorithm, includes:
[0036] Step S521, deleting the stop words of each word segmentation sub-data in the current data set to be processed;
[0037] Step S522, respectively obtaining the occurrence frequency of each word segmentation sub-data in the current data set to be processed;
[0038] Step S523, respectively obtaining the inverse text frequency of each word segmentation sub-data based on all the data sets to be processed;
[0039] Step S524, respectively obtaining the product of the occurrence frequency of each word segmentation sub-data and the inverse text frequency;
[0040] Step S525, sorting the word segmentation sub-data of the current data set to be processed from large to small based on the values of all products, to obtain a sorting table of the word segmentation sub-data of the current data set to be processed;
[0041] Step S526: Obtain one or more word segmentation sub-data that are ranked top in the current word segmentation sub-data ranking table and define them as keywords.
[0042] As a further improvement of the present application, step S6, classifies all keywords according to the sentiment polarity and the occurrence timestamp, and obtains a sentiment polarity classification table based on a preset time period, including:
[0043] Step S61, obtaining all keywords within the current preset time period and defining the keyword data set of the current preset time period , For the word segmentation data, is the number of all word segmentation sub-data;
[0044] Step S62: defining a set of emotion polarity categories according to a plurality of preset types of emotion polarity ,in, For the category set The Emotional polarity, is the number of types of all sentiment polarities;
[0045] Step S63, calculate the first The conditional probability of each word segmentation data under each sentiment polarity:
[0046] (1);
[0047] in, For the The conditional probability of keyword dataset under sentiment polarity; For the The marginal probability of a sentiment polarity; For the Emotional polarity The conditional probability of word segmentation data;
[0048] Step S64, classifying each word segmentation sub-data into the sentiment polarity with the highest conditional probability;
[0049] Step S65, integrating all classified word segmentation sub-data to obtain a sentiment polarity classification table for the current preset time period.
[0050] As a further improvement of the present application, step S9, through the SQL query statement, SQL update statement, SQL delete statement of the MySQL client, the data query, data update, data deletion of the tin indium material public opinion database is performed, and then, it includes:
[0051] Step S100, generating a query record with a query timestamp and query content based on each data query;
[0052] Step S200, generating an update record with an update timestamp and update content based on each data update;
[0053] Step S300, generating a deletion record with a deletion timestamp and deletion content based on each data deletion;
[0054] Step S400, sending all query records, all update records, and all deletion records to an external monitoring terminal.
[0055] In order to achieve the above objectives, this application also provides the following technical solutions:
[0056] A data management device for public opinion on tin and indium materials, the data management device is applied to the above-mentioned data management method, and the data management device comprises:
[0057] The global environment variable file building module is used to build the global environment variable file of TiDB through the TiUP component manager;
[0058] A TiDB cluster generation module, used to generate a TiDB cluster in the global environment variable file through the playground instruction of the TiUP component manager;
[0059] A TiDB cluster connection module, used to connect to the TiDB cluster through a MySQL client signal;
[0060] A tin-indium material public data crawling module, used to crawl a number of public data related to tin-indium materials on the public webpage through a data crawler script;
[0061] The keyword extraction module is used to extract features from all public data and obtain several keywords;
[0062] The sentiment polarity classification table acquisition module is used to classify all keywords according to the sentiment polarity and occurrence timestamp, and obtain a sentiment polarity classification table based on a preset time period;
[0063] A blank database creation module, used to create a blank database with the same dimension as the sentiment polarity classification table in the TiDB cluster through the SQL creation statement of the MySQL client;
[0064] A tin indium material public opinion database generation module is used to import all sentiment polarity classification tables into the blank database through the SQL import statement of the MySQL client to form a tin indium material public opinion database;
[0065] The tin and indium material public opinion database management module is used to query, update and delete data in the tin and indium material public opinion database through the SQL query statement, SQL update statement and SQL delete statement of the MySQL client.
[0066] In order to achieve the above objectives, this application also provides the following technical solutions:
[0067] An electronic device comprises a processor and a memory coupled to the processor, wherein the memory stores program instructions executable by the processor; when the processor executes the program instructions stored in the memory, the data management method as described above is implemented.
[0068] In order to achieve the above objectives, this application also provides the following technical solutions:
[0069] A storage medium stores program instructions, and when the program instructions are executed by a processor, the data management method as described above can be implemented.
[0070] This application builds a global environment variable file of TiDB through the TiUP component manager; generates a TiDB cluster in the global environment variable file through the playground instruction of the TiUP component manager; connects the TiDB cluster through a MySQL client signal; crawls a number of public data related to tin and indium materials on public web pages through a data crawler script; extracts features from all public data to obtain a number of keywords; classifies all keywords according to sentiment polarity and occurrence timestamp, and obtains a sentiment polarity classification table based on a preset time period; creates a blank database with the same dimension as the sentiment polarity classification table in the TiDB cluster through the SQL creation statement of the MySQL client; imports all sentiment polarity classification tables into the blank database through the SQL import statement of the MySQL client to form a tin and indium material public opinion database; and performs data query, data update, and data deletion on the tin and indium material public opinion database through the SQL query statement, SQL update statement, and SQL delete statement of the MySQL client. This application takes advantage of the rapid deployment feature of TiDB, and quickly completes the installation and configuration of TiDB through simple command lines or graphical interface tools, which significantly shortens the time from environment preparation to database operation and improves management efficiency. At the same time, this application also takes advantage of the good compatibility of TiDB and MySQL, and management operations can be performed through SQL statements. The database is constructed in a form with the same dimensionality as the classified data, so that the classified public opinion data can be directly loaded into the database and directly retrieved without the need for additional conversion operations. The TiUP component manager can guide users to build databases quickly and efficiently, further improving management efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Figure 1 This is a schematic diagram of the steps of an embodiment of a method for data management of public opinion on tin and indium materials in the present application;
[0072] Figure 2 This is a functional module diagram of an embodiment of a data management device for public opinion on tin and indium materials of the present application;
[0073] Figure 3 This is a schematic diagram of the structure of an embodiment of the electronic device of the present application;
[0074] Figure 4 This is a schematic diagram of the structure of an embodiment of the storage medium of the present application. DETAILED DESCRIPTION
[0075] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0076] The terms "first", "second" and "third" in this application are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined as "first", "second" and "third" can explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "multiple" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined. All directional indications (such as up, down, left, right, front, back...) in the embodiments of this application are only used to explain the relative position relationship, movement, etc. between the components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication also changes accordingly. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally also includes steps or units that are not listed, or optionally also includes other steps or units inherent to these processes, methods, products or devices.
[0077] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0078] like Figure 1 As shown, this embodiment provides an embodiment of a method for data management of public opinion on tin-indium materials. In this embodiment, public opinion on tin-indium materials appears on a public webpage of the alloy material industry.
[0079] For example, public web pages include news pages, forum pages, post bar pages, blog pages, etc., as well as portal pages of various exhibitions.
[0080] Preferably, the public opinion of this embodiment can be the technical public opinion of tin indium materials, such as the current trend of technology, technical advantages and disadvantages, and technology application prospects; it can also be the social public opinion of tin indium materials, such as usage experience, perception, and subjective evaluation; it can even be the surrounding public opinion of tin indium materials, such as exhibition conditions, exhibition exposure, and exhibition evaluation.
[0081] Specifically, the data management method includes the following steps:
[0082] Step S1: Build the global environment variable file (TiUP Playground) of TiDB through the TiUP component manager.
[0083] Glossary: TiDB (Ti Distributed Database) is an open source distributed relational database. TiDB is a hybrid transactional and analytical processing (HTAP) distributed database that supports both online transaction processing (OLTP) and online analytical processing (OLAP). It has important features such as horizontal expansion or contraction, financial-grade high availability, real-time HTAP, cloud native, and is compatible with the MySQL 5.7 protocol and ecosystem. TiDB is suitable for various scenarios with high availability and strong consistency requirements, especially in environments with large data scale and high concurrency requirements, combining the best features of traditional RDBMS and NoSQL. TiDB is compatible with MySQL and supports unlimited horizontal expansion, which is conducive to the data expansion of this embodiment, and has strong consistency and high availability.
[0084] Step S2: Generate a TiDB cluster in the global environment variable file through the playground instruction of the TiUP component manager.
[0085] Preferably, the most basic TiDB test cluster usually consists of 2 TiDB instances, 3 TiKV instances, 3 PD instances, and an optional TiFlash instance. Through TiUP Playground, you can quickly build a basic test cluster.
[0086] Preferably, the above TiDB instance, TiKV instance, PD instance, and TiFlash instance can be implemented through the following commands:
[0087] [root@worker2 tidb]# tiup playground v7.5.0 --db 2 --pd 3 --kv 3.
[0088] It is worth noting that after TiUP is installed, it will prompt the absolute path of the Shell profile file. Before executing the source command ([root@worker2 tidb]# source / root / .bash_profile), you need to modify ${your_shell_profile} to the actual location of the Shell profile file.
[0089] Step S3: Connect to the TiDB cluster through the MySQL client signal.
[0090] Preferably, you can use the following command to connect the MySQL client to the TiDB cluster:
[0091] [root@worker2 tidb]# mysql --host 127.0.0.1 --port 4000 -u root.
[0092] Step S4, crawling a number of public data related to tin and indium materials from public web pages through a data crawler script.
[0093] Preferably, the data crawler script can be based on the Scrapy framework. The crawler built by the framework is written in Python and can quickly traverse website data according to user needs. Unlike traditional crawler programs, the Scrapy crawler can also crawl the website's API data interface, thereby increasing the speed of crawling information.
[0094] Preferably, the crawler structure based on the Scrapy framework includes a crawler script body, a crawler engine, a scheduling plug-in, a download module, a crawler middleware and a pipeline. The target of the crawler script body is the URL address, and the crawler sends the content of the target URL address into the pipeline for storage; the crawler engine is responsible for transmitting content data in all modules; the scheduling plug-in is to schedule the resource requests required by the engine; the download module is controlled by the crawler script, and when the crawler needs to download web page content, it will call the downloader to download.
[0095] Step S5: extract features from all public data to obtain several keywords.
[0096] Step S6, classify all keywords according to the sentiment polarity and occurrence timestamp, and obtain a sentiment polarity classification table based on a preset time period.
[0097] For example: The sentiment polarity classification table (from XX / XX / XXXX to XX / XX / XY / XXXX) is shown in Table 1 below:
[0098]
[0099] Table 1: Sentiment polarity classification table (from XX / XX / XXXX to XX / XX / XXXX).
[0100] It is worth noting that Table 1 above is obtained based on the classification of tin and indium related public opinion data on a natural day, that is, a Table 1 will be obtained for each natural day, and the data in Table 1 on each natural day may be the same or different.
[0101] It is worth noting that the data in Table 1 above are only for illustrative purposes and do not represent the actual situation.
[0102] Step S7: Create a blank database with the same dimension as the sentiment polarity classification table in the TiDB cluster through the SQL creation statement of the MySQL client.
[0103] Preferably, the SQL creation statement is:
[0104] mysql> CREATE DATABASE your_tidb.
[0105] Preferably, the dimension of the above Table 1 is 4×n, and the dimension of the blank database is also set to 4×n.
[0106] It is worth noting that the dimension of the blank database is only for illustration and does not represent the actual situation.
[0107] Step S8, import all sentiment polarity classification tables into a blank database through the SQL import statement of the MySQL client to form a tin and indium material public opinion database.
[0108] Preferably, the SQL import statement is:
[0109] mysql> INSERT INTO your_tidb VALUES(1,2,3,4); The statement here means to add 1,2,3,4 in the data (1,2,3,4) to the table in the order of the four headers of Table 1.
[0110] Step S9, query, update and delete data on the tin and indium material public opinion database through the SQL query statement, SQL update statement and SQL delete statement of the MySQL client.
[0111] Preferably, the SQL query statement is:
[0112] mysql> SELECT * FROM your_tidb.
[0113] mysql> SELECT * FROM your_tidb WHERE id<5; The statement here means to query the data with a sequence number less than 5.
[0114] The SQL update statement is:
[0115] mysql> UPDATE your_tidb SET (3)='400' WHERE id=6; The statement here means to replace the data with serial number 6 in the third column with "400", that is, to change the 460 in the conductivity row to 400.
[0116] The SQL delete statement is:
[0117] mysql> DELETE FROM your_tidb WHERE id=2; The statement here means to delete the entire row of data with sequence number 2.
[0118] It should be noted that the above statements are only examples and are not used to limit the fixed format of the statements. Other SQL call statements can also directly call the database of this embodiment.
[0119] Further, step S8, importing all sentiment polarity classification tables into a blank database through the SQL import statement of the MySQL client to form a tin and indium material public opinion database, and then including:
[0120] Step S10, splitting the tin and indium material public opinion database into a plurality of sub-databases based on a preset time period.
[0121] Step S20, obtain the on-chain hash certificate of each sub-database through the SHA256 algorithm.
[0122] Preferably, the SHA256 algorithm is an algorithm subdivided from SHA-2. The name SHA-2 comes from the abbreviation of Secure Hash Algorithm 2 (English: Secure Hash Algorithm 2), which is a cryptographic hash function algorithm standard developed by the National Security Agency of the United States. It is one of the SHA algorithms and the successor of SHA-1.
[0123] SHA-2 can be further divided into six different algorithm standards: SHA224, SHA256, SHA384, SHA512, SHA512 / 224, SHA512 / 256. These variants have the same basic structure except for some minor differences such as the length of the generated summary and the number of loop runs. SHA-256 is a hash function, also known as a hash algorithm, which is a method of creating a small digital "fingerprint" from any kind of data. The hash function compresses the message or data into a summary, making the data smaller and fixing the format of the data. The function shuffles and mixes the data to recreate a fingerprint called a hash value (or hash value). The hash value is usually represented by a short string of random letters and numbers. For any length of message, SHA256 will generate a 256-bit hash value, called a message digest, which is equivalent to an array of 32 bytes in length, usually represented by a hexadecimal string of 64 in length.
[0124] Step S30, copy all the on-chain hash evidence and save them locally, forming a local hash evidence based on a sub-database.
[0125] Preferably, on-chain hash evidence and local hash evidence can ensure the authenticity of the data.
[0126] Step S40, save all the hash evidence on the chain to the blockchain in order from old to new according to the preset time period.
[0127] Preferably, the above-mentioned natural day can be used as a storage node, and the chain span of the natural day can be used to realize the chain span of the on-chain hash evidence.
[0128] Step S50, based on the current local hash evidence, retrieve the on-chain hash evidence that matches the local hash evidence in the blockchain.
[0129] Step S60, determine whether the on-chain hash evidence is retrieved successfully. If the on-chain hash evidence is not retrieved successfully, execute step S70.
[0130] Step S70, determine whether one of the on-chain hash evidence and the local hash evidence has been tampered with.
[0131] Step S80: Generate a data tampering signal and send it to an external monitoring terminal.
[0132] Preferably, hash evidence is to save the hash value of the file content on the chain. The hash value of the file content is usually called the "digital fingerprint" of the file, which can be obtained by hashing the file content. Because the length of the hash value is relatively limited, for example, the SHA256 hash value of a content of tens of thousands of words is only 256 bits of characters, and it is no pressure for the blockchain to store such a length of content on the chain. The hash evidence can be used to verify whether the file content has been tampered with. For example, the hash value of an original text is stored on the blockchain. When we get this file again, we perform a hash operation on its content. If it is consistent with the content stored on the chain, it is considered that the content is credible and has not been tampered with. If the hash value is different, it is considered that the content has been tampered with and is no longer credible. To prevent the data from being maliciously implanted with viruses, the company can put the "digital fingerprint" of the public opinion data it has obtained into the blockchain. Users can verify whether the digital fingerprint has changed when downloading software from different channels. If there is a change, it is considered that the software may have been implanted with a virus and is no longer safe.
[0133] Furthermore, in step S5, all public data are subjected to feature extraction to obtain several keywords, including:
[0134] Step S51, performing word segmentation processing on each public data by using a word segmentation algorithm, and obtaining a data set to be processed having a plurality of word segmentation sub-data based on one public data.
[0135] Preferably, the word segmentation algorithm is the basic technology in natural language processing. Its task is to segment a continuous text sequence into meaningful word units, that is, vocabulary units. Common word segmentation algorithms are as follows:
[0136] ① Rule-based word segmentation: word segmentation is performed through a predefined rule library, such as regular expressions, dictionary matching, etc., which is suitable for some highly structured languages.
[0137] ② Statistical word segmentation: Utilize the learning of large-scale corpus, such as Hidden Markov Model (HMM) and Conditional Random Field (CRF), to segment words according to contextual probability.
[0138] ③ Word segmentation based on word frequency statistics: According to the word frequency table, words with higher frequency are given priority for segmentation, which is suitable for simple texts.
[0139] ④ Deep learning word segmentation: Use neural network models, such as recurrent neural network (RNN), long short-term memory network (LSTM) or Transformer, to perform word segmentation through end-to-end learning.
[0140] ⑤ Adaptive Dynamic Vocabulary (ADV): It combines dictionary and statistical information to handle both common words and new words.
[0141] Step S52: extract keywords of all word segmentation sub-data in the current data set to be processed by using a feature extraction algorithm.
[0142] Preferably, the feature extraction algorithm can adopt the TF-IDF algorithm, which means term frequency-inverse text frequency. The TF in the algorithm is term frequency, which is usually used to measure the frequency of a word appearing in the entire text. The IDF in the algorithm is the inverse text frequency, that is, the reciprocal of the number of times a word appears in a text. The algorithm can indicate the importance of a word in a text, that is, a keyword.
[0143] Furthermore, in step S51, each public data is segmented by a segmentation algorithm, and a data set to be processed having a plurality of segmentation sub-data is obtained based on the public data, including:
[0144] Step S511, performing word segmentation processing on each public data through SentencePiece to mark the word segmentation point of each public data.
[0145] Preferably, in addition to the SentencePiece software, Stanford NLP and moses tokenizer tools can also be used to perform word segmentation and byte pair processing operations on the data.
[0146] Step S512, dividing the current public data into a number of segmentations based on all segmentation points.
[0147] Step S513: Mark each word segmentation by decoding the time tag, and obtain a word segmentation sub-data based on one mark.
[0148] Specifically, the decoding time tag is a DTS tag, and PAD, UNK, and EOS tags may also be used.
[0149] Step S514: Pack all the word segmentation sub-data of the current public data into a data set to be processed.
[0150] Furthermore, in step S52, keywords of all word segmentation sub-data in the current data set to be processed are extracted by a feature extraction algorithm, including:
[0151] Step S521, deleting the stop words of each word segmentation sub-data in the current data set to be processed.
[0152] Preferably, the stop words are meaningless words such as punctuation marks, auxiliary words, conjunctions, etc., such as "de", "le", "a", etc., and the Python stop word list can be directly used.
[0153] Step S522: Obtain the occurrence frequency of each segmented sub-data in the current dataset to be processed respectively.
[0154] Preferably, the occurrence frequency is equal to the number of occurrences of the segmented sub-data in the current dataset to be processed divided by the total number of words in the current dataset to be processed.
[0155] Step S523: Obtain the inverse document frequency of each segmented sub-data based on all datasets to be processed respectively.
[0156] Preferably, the inverse document frequency is equal to the logarithm to the base 10 of the result obtained by dividing all datasets to be processed by the number of documents containing the segmented sub-data plus 1.
[0157] Step S524: Obtain the product of the occurrence frequency and the inverse document frequency of each segmented sub-data respectively.
[0158] Step S525: Sort the segmented sub-data of the current dataset to be processed in descending order based on the values of all products, and obtain the sorting table of the segmented sub-data of the current dataset to be processed.
[0159] Step S526: Obtain one or more segmented sub-data ranked靠前 in the current sorting table of segmented sub-data and define them as keywords.
[0160] For example, the sorting table 2 (in reverse order) of the segmented sub-data of the current dataset to be processed is as follows:
[0161]
[0162] Table 2: Sorting table (in reverse order) of the segmented sub-data of the current dataset to be processed.
[0163] Preferably, it is only necessary to sort the segmented sub-data of the current dataset to be processed in descending order based on the values of all products. In this embodiment, the reverse order is shown to illustrate that the product is directly proportional to the inverse document frequency of a word and inversely proportional to the number of occurrences of the word. Therefore, keywords can be extracted by calculating the product value of each word, then sorting them in descending order, and taking the first few words.
[0164] It should be noted that the words in Table 2 above are only for principle illustration and do not represent the real situation.
[0165] Further, in step S6, all keywords are classified according to the sentiment polarity and the occurrence timestamp, and a sentiment polarity classification table is obtained based on a preset time period, including:
[0166] Step S61, obtaining all keywords within the current preset time period and defining the keyword data set of the current preset time period , For the word segmentation data, is the number of all word segmentation sub-data.
[0167] Step S62: defining a set of emotion polarity categories according to a plurality of preset types of emotion polarity ,in, For category collection The Emotional polarity, is the number of types of all sentiment polarities.
[0168] Step S63, calculate the first The conditional probability of each word segmentation data under each sentiment polarity:
[0169] (1).
[0170] in, For the The conditional probability of keyword dataset under sentiment polarity; For the The marginal probability of a sentiment polarity; For the Emotional polarity The conditional probability of the word segmentation data.
[0171] Step S64: classify each word segmentation sub-data into the sentiment polarity with the highest conditional probability.
[0172] Step S65, integrating all classified word segmentation sub-data to obtain a sentiment polarity classification table for the current preset time period.
[0173] Preferably, the naive Bayes classification assumes that the existence of a specific feature in a class is independent of the existence of any other feature, that is, each feature is independent of each other. Therefore, there are some constraints on the actual situation. If there is a correlation between attributes, the classification accuracy will be reduced, but in actual application, the classification effect of naive Bayes is relatively accurate. Naive Bayes solves the probability of each category appearing under the condition that the item appears for a given item to be classified. The largest probability is considered to be the category to which the item to be classified belongs.
[0174] Specifically, Naive Bayes is defined as follows:
[0175] ① Let x={a1, a2, a3,…, an} be an item to be classified, and each a is a feature of x.
[0176] ②There is a category set c={y1, y2, y3,…, ym}.
[0177] ③Calculate P(y1|x), P(y2|x), …, P(ym|x).
[0178] ④If P(yk|x)=max{P(y1|x),P(y2|x),…,P(ym|x)}, then x∈yk.
[0179] Then calculate the conditional probabilities in step ③ through the following steps:
[0180] Find a set of items to be classified with known classifications. This set is called a training sample set.
[0181] The conditional probability estimates of each feature attribute in each category are obtained by statistics. That is:
[0182] P(a1|y1),P(a2|y1),……,P(an|y1)
[0183] P(a1|y2),P(a2|y2),…,P(an|y2);
[0184] …
[0185] P(a1|ym),P(a2|ym),…,P(an|ym);
[0186] Assuming that each feature attribute is conditionally independent, according to the Bayesian principle:
[0187] P(yi|x)=P(x|yi)P(yi) / p(x).
[0188] Since the denominator is a constant for all categories, we only need to maximize the numerator. And because each feature attribute is conditionally independent, then:
[0189] P(x|yi)P(yi)=P(a1|yi)P(a2|yi)…P(an|yi)P(yi).
[0190] It should be noted that the above preferred contents are for explanation of the principle, and the meaning of their symbols is not interchangeable with the meaning of the symbols of other formulas in this embodiment.
[0191] Further, step S9, through the SQL query statement, SQL update statement, SQL delete statement of the MySQL client, the tin indium material public opinion database is queried, updated, and deleted, and then includes:
[0192] Step S100, generating a query record with a query timestamp and query content based on each data query.
[0193] Step S200: Generate an update record with an update timestamp and update content based on each data update.
[0194] Step S300: Generate a deletion record with a deletion timestamp and deletion content based on each data deletion.
[0195] Step S400, sending all query records, all update records, and all deletion records to an external monitoring terminal.
[0196] In this embodiment, a global environment variable file of TiDB is constructed through the TiUP component manager; a TiDB cluster is generated in the global environment variable file through the playground instruction of the TiUP component manager; a TiDB cluster is connected through a MySQL client signal; a number of public data related to tin and indium materials are crawled from public web pages through a data crawler script; features are extracted from all public data to obtain a number of keywords; all keywords are classified according to sentiment polarity and occurrence timestamp, and a sentiment polarity classification table is obtained based on a preset time period; a blank database with the same dimension as the sentiment polarity classification table is created in the TiDB cluster through the SQL creation statement of the MySQL client; all sentiment polarity classification tables are imported into the blank database through the SQL import statement of the MySQL client to form a tin and indium material public opinion database; data query, data update, and data deletion are performed on the tin and indium material public opinion database through the SQL query statement, SQL update statement, and SQL delete statement of the MySQL client. This embodiment takes advantage of the rapid deployment feature of TiDB, and quickly completes the installation and configuration of TiDB through a simple command line or graphical interface tool, which significantly shortens the time from environment preparation to database operation and improves management efficiency. At the same time, this embodiment also takes advantage of the good compatibility of TiDB and MySQL, and management operations can be performed through SQL statements. The database is constructed in a form with the same dimensionality as the classified data, so that the classified public opinion data can be directly loaded into the database and directly retrieved without the need for additional conversion operations. The TiUP component manager can guide users to build databases quickly and efficiently, further improving management efficiency.
[0197] like Figure 2 As shown, this embodiment provides an embodiment of a data management device for public opinion on tin and indium materials. In this embodiment, the data management device is applied to the data management method in the above embodiment.
[0198] Specifically, the data management device includes a global environment variable file construction module 1, a TiDB cluster generation module 2, a TiDB cluster connection module 3, a tin indium material public data crawling module 4, a keyword extraction module 5, a sentiment polarity classification table acquisition module 6, a blank database creation module 7, a tin indium material public opinion database generation module 8, and a tin indium material public opinion database management module 9, which are electrically connected in sequence.
[0199] Among them, the global environment variable file construction module 1 is used to build the global environment variable file of TiDB through the TiUP component manager; the TiDB cluster generation module 2 is used to generate a TiDB cluster in the global environment variable file through the playground instruction of the TiUP component manager; the TiDB cluster connection module 3 is used to connect to the TiDB cluster through the MySQL client signal; the tin indium material public data crawling module 4 is used to crawl a number of public data related to tin indium materials on public web pages through the data crawler script; the keyword extraction module 5 is used to extract features from all public data to obtain a number of keywords; the sentiment polarity classification table acquisition module 6 is used to classify the sentiment polarity and the occurrence timestamp All keywords are classified and processed, and a sentiment polarity classification table is obtained based on a preset time period; the blank database creation module 7 is used to create a blank database with the same dimension as the sentiment polarity classification table in the TiDB cluster through the SQL creation statement of the MySQL client; the tin indium material public opinion database generation module 8 is used to import all sentiment polarity classification tables into the blank database through the SQL import statement of the MySQL client to form a tin indium material public opinion database; the tin indium material public opinion database management module 9 is used to query, update and delete data in the tin indium material public opinion database through the SQL query statement, SQL update statement and SQL delete statement of the MySQL client.
[0200] Furthermore, the data management device also includes a tin-indium material public opinion database splitting module, an on-chain hash evidence acquisition module, a local hash evidence acquisition module, an on-chain hash evidence on-chain module, a hash evidence retrieval module, a hash evidence verification module, a hash evidence determination module, and a data tampering signal generation and sending module, which are electrically connected in sequence; wherein the tin-indium material public opinion database splitting module is electrically connected to the tin-indium material public opinion database generation module 8.
[0201] Among them, the tin and indium material public opinion database splitting module is used to split the tin and indium material public opinion database into several sub-databases based on a preset time period; the on-chain hash evidence acquisition module is used to obtain the on-chain hash evidence of each sub-database through the SHA256 algorithm; the local hash evidence acquisition module is used to copy all on-chain hash evidence and save them locally, forming a local hash evidence based on a sub-database; the on-chain hash evidence on-chain module is used to save all on-chain hash evidence in sequence from old to new according to the preset time period to the blockchain; the hash evidence retrieval module is used to retrieve the on-chain hash evidence that matches the local hash evidence in the blockchain based on the current local hash evidence; the hash evidence verification module is used to determine whether the on-chain hash evidence has been retrieved successfully; the hash evidence judgment module is used to determine that one of the on-chain hash evidence and the local hash evidence has been tampered with if the on-chain hash evidence has not been retrieved successfully; the data tampering signal generation and sending module is used to generate a data tampering signal and send it to an external monitoring terminal.
[0202] Furthermore, the keyword extraction module 5 specifically includes a first keyword extraction submodule and a second keyword extraction submodule which are electrically connected in sequence; the first keyword extraction submodule is electrically connected to the tin-indium material public data crawling module 4, and the second keyword extraction submodule is electrically connected to the sentiment polarity classification table acquisition module 6.
[0203] Among them, the first keyword extraction submodule is used to perform word segmentation processing on each public data through a word segmentation algorithm, and obtain a data set to be processed with several word segmentation sub-data based on one public data; the second keyword extraction submodule is used to extract keywords of all word segmentation sub-data in the current data set to be processed through a feature extraction algorithm.
[0204] Furthermore, the first keyword extraction submodule specifically includes a first keyword extraction unit, a second keyword extraction unit, a third keyword extraction unit, and a fourth keyword extraction unit which are electrically connected in sequence; the first keyword extraction unit is electrically connected to the tin indium material public data crawling module 4, and the fourth keyword extraction unit is electrically connected to the second keyword extraction submodule.
[0205] Among them, the first keyword extraction unit is used to perform word segmentation processing on each public data through SentencePiece to mark the word segmentation point of each public data; the second keyword extraction unit is used to divide the current public data into several word segments based on all word segmentation points; the third keyword extraction unit is used to mark each word segmentation by decoding the time tag, and obtain a word segmentation sub-data based on a tag; the fourth keyword extraction unit is used to package all the word segmentation sub-data of the current public data into a data set to be processed.
[0206] Furthermore, the second keyword extraction submodule specifically includes a fifth keyword extraction unit, a sixth keyword extraction unit, a seventh keyword extraction unit, an eighth keyword extraction unit, a ninth keyword extraction unit, and a tenth keyword extraction unit which are electrically connected in sequence; the fifth keyword extraction unit is electrically connected to the fourth keyword extraction unit, and the tenth keyword extraction unit is electrically connected to the sentiment polarity classification table acquisition module 6.
[0207] Among them, the fifth keyword extraction unit is used to delete the stop words of each segmentation sub-data in the current data set to be processed; the sixth keyword extraction unit is used to respectively obtain the occurrence frequency of each segmentation sub-data in the current data set to be processed; the seventh keyword extraction unit is used to respectively obtain the inverse text frequency of each segmentation sub-data based on all data sets to be processed; the eighth keyword extraction unit is used to respectively obtain the product of the occurrence frequency and the inverse text frequency of each segmentation sub-data; the ninth keyword extraction unit is used to sort the segmentation sub-data of the current data set to be processed from large to small based on the values of all products, and obtain the segmentation sub-data sorting table of the current data set to be processed; the tenth keyword extraction unit is used to obtain one or more segmentation sub-data ranked at the top in the current segmentation sub-data sorting table and define them as keywords.
[0208] Furthermore, the emotion polarity classification table acquisition module 6 specifically includes a first emotion polarity classification table acquisition submodule, a second emotion polarity classification table acquisition submodule, a third emotion polarity classification table acquisition submodule, a fourth emotion polarity classification table acquisition submodule, and a fifth emotion polarity classification table acquisition submodule, which are electrically connected in sequence; the first emotion polarity classification table acquisition submodule is electrically connected to the tenth keyword extraction unit, and the fifth emotion polarity classification table acquisition submodule is electrically connected to the blank database creation module 7.
[0209] The first sentiment polarity classification table acquisition submodule is used to acquire all keywords within the current preset time period and define the keyword data set of the current preset time period. , For the word segmentation data, is the number of all word segmentation sub-data.
[0210] The second emotion polarity classification table acquisition submodule is used to define an emotion polarity category set according to several preset types of emotion polarity. ,in, For category collection The Emotional polarity, is the number of types of all sentiment polarities.
[0211] The third sentiment polarity classification table acquisition submodule is used to calculate the first The conditional probability of each word segmentation data under each sentiment polarity:
[0212] (1).
[0213] in, For the The conditional probability of keyword dataset under sentiment polarity; For the The marginal probability of a sentiment polarity; For the Emotional polarity The conditional probability of the word segmentation data.
[0214] The fourth sentiment polarity classification table acquisition submodule is used to classify each word segmentation sub-data into the sentiment polarity with the highest conditional probability.
[0215] The fifth sentiment polarity classification table acquisition submodule is used to integrate all classified word segmentation sub-data to obtain the sentiment polarity classification table of the current preset time period.
[0216] Furthermore, the data management device also includes a query record generation module, an update record generation module, a deletion record generation module, and an operation record sending module which are electrically connected in sequence; the query record generation module is electrically connected to the tin-indium material public opinion database management module 9.
[0217] Among them, the query record generation module is used to generate a query record with a query timestamp and query content based on each data query; the update record generation module is used to generate an update record with an update timestamp and update content based on each data update; the deletion record generation module is used to generate a deletion record with a deletion timestamp and deletion content based on each data deletion; the operation record sending module is used to send all query records, all update records, and all deletion records to the external monitoring terminal.
[0218] It should be noted that this embodiment is a functional module embodiment based on the above method embodiment. The optimization, expansion, limitation, example, and principle description of this embodiment can be referred to the above embodiment, and this embodiment will not be repeated.
[0219] In this embodiment, a global environment variable file of TiDB is constructed through the TiUP component manager; a TiDB cluster is generated in the global environment variable file through the playground instruction of the TiUP component manager; a TiDB cluster is connected through a MySQL client signal; a number of public data related to tin and indium materials are crawled from public web pages through a data crawler script; features are extracted from all public data to obtain a number of keywords; all keywords are classified according to sentiment polarity and occurrence timestamp, and a sentiment polarity classification table is obtained based on a preset time period; a blank database with the same dimension as the sentiment polarity classification table is created in the TiDB cluster through the SQL creation statement of the MySQL client; all sentiment polarity classification tables are imported into the blank database through the SQL import statement of the MySQL client to form a tin and indium material public opinion database; data query, data update, and data deletion are performed on the tin and indium material public opinion database through the SQL query statement, SQL update statement, and SQL delete statement of the MySQL client. This embodiment takes advantage of the rapid deployment feature of TiDB, and quickly completes the installation and configuration of TiDB through a simple command line or graphical interface tool, which significantly shortens the time from environment preparation to database operation and improves management efficiency. At the same time, this embodiment also takes advantage of the good compatibility of TiDB and MySQL, and management operations can be performed through SQL statements. The database is constructed in a form with the same dimensionality as the classified data, so that the classified public opinion data can be directly loaded into the database and directly retrieved without the need for additional conversion operations. The TiUP component manager can guide users to build databases quickly and efficiently, further improving management efficiency.
[0220] Figure 3 is a schematic diagram of the structure of an electronic device according to an embodiment of the present application. Figure 3 As shown, the electronic device 10 includes a processor 101 and a memory 102 coupled to the processor 101 .
[0221] The memory 102 stores program instructions for implementing a data management method for public opinion on tin and indium materials in any of the above embodiments.
[0222] The processor 101 is used to execute program instructions stored in the memory 102 to perform data management of public opinion on tin and indium materials.
[0223] The processor 101 may also be referred to as a CPU (Central Processing Unit). The processor 101 may be an integrated circuit chip having a signal processing capability. The processor 101 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0224] Further, Figure 4 This is a schematic diagram of the structure of a storage medium according to an embodiment of the present application. Figure 4 As shown, the storage medium 11 of the embodiment of the present application stores program instructions 111 that can implement all the above methods, wherein the program instructions 111 can be stored in the above storage medium in the form of a software product, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, or terminal devices such as a computer, a server, a mobile phone, and a tablet.
[0225] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0226] In addition, each functional unit in each embodiment of the present application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of software functional units. The above is only an implementation method of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the specification and drawings of this application, or directly or indirectly used in other related technical fields, is also included in the patent protection scope of the present application.
Claims
1. A data management method for public opinion on tin and indium materials, wherein the public opinion on tin and indium materials appears on a public webpage of the alloy material industry, characterized in that: The data management method comprises: Step S1, build the global environment variable file of TiDB through the TiUP component manager; Step S2, generating a TiDB cluster in the global environment variable file through the playground instruction of the TiUP component manager; Step S3, connecting to the TiDB cluster via a MySQL client signal; Step S4, crawling a number of public data related to tin and indium materials from the public webpage through a data crawler script; Step S5, extracting features from all public data to obtain several keywords; Step S6, classifying all keywords according to the sentiment polarity and occurrence timestamp, and obtaining a sentiment polarity classification table based on a preset time period; Step S7, creating a blank database with the same dimension as the sentiment polarity classification table in the TiDB cluster through the SQL creation statement of the MySQL client; Step S8, importing all sentiment polarity classification tables into the blank database through the SQL import statement of the MySQL client to form a tin and indium material public opinion database; Step S9, query, update and delete data on the tin-indium material public opinion database through the SQL query statement, SQL update statement and SQL delete statement of the MySQL client.
2. The data management method according to claim 1, characterized in that: Step S8, importing all sentiment polarity classification tables into the blank database through the SQL import statement of the MySQL client to form a tin-indium material public opinion database, and then comprising: Step S10, splitting the tin-indium material public opinion database into a plurality of sub-databases based on the preset time period; Step S20, obtaining the chain hash certificate of each sub-database through the SHA256 algorithm; Step S30, copy all the hash certificates on the chain and save them locally, so as to obtain a local hash certificate based on a sub-database; Step S40, saving all the hash evidence on the blockchain in order from old to new according to the preset time period; Step S50, based on the current local hash evidence, retrieve the on-chain hash evidence that matches the local hash evidence in the blockchain; Step S60, determining whether the on-chain hash evidence is retrieved successfully, if the on-chain hash evidence is not retrieved successfully, executing step S70; Step S70, determining whether one of the on-chain hash evidence and the local hash evidence has been tampered with; Step S80: Generate a data tampering signal and send it to an external monitoring terminal.
3. The data management method according to claim 1, characterized in that: Step S5: extract features from all public data to obtain several keywords, including: Step S51, performing word segmentation processing on each public data by using a word segmentation algorithm, and obtaining a data set to be processed having a plurality of word segmentation sub-data based on one public data; Step S52: extract keywords of all word segmentation sub-data in the current data set to be processed by using a feature extraction algorithm.
4. The data management method according to claim 3, characterized in that: Step S51, perform word segmentation processing on each public data by using a word segmentation algorithm, and obtain a data set to be processed having a plurality of word segmentation sub-data based on one public data, including: Step S511, performing word segmentation processing on each public data by using SentencePiece to mark the word segmentation point of each public data; Step S512, dividing the current public data into a number of segmentations based on all segmentation points; Step S513, marking each word segmentation by decoding the time tag, and obtaining a word segmentation sub-data based on one mark; Step S514: Pack all the word segmentation sub-data of the current public data into a data set to be processed.
5. The data management method according to claim 3, characterized in that: Step S52, extracting keywords of all word segmentation sub-data in the current data set to be processed by a feature extraction algorithm, including: Step S521, deleting the stop words of each word segmentation sub-data in the current data set to be processed; Step S522, respectively obtaining the occurrence frequency of each word segmentation sub-data in the current data set to be processed; Step S523, respectively obtaining the inverse text frequency of each word segmentation sub-data based on all the data sets to be processed; Step S524, respectively obtaining the product of the occurrence frequency of each word segmentation sub-data and the inverse text frequency; Step S525, sorting the word segmentation sub-data of the current data set to be processed from large to small based on the values of all products, to obtain a sorting table of the word segmentation sub-data of the current data set to be processed; Step S526: Obtain one or more word segmentation sub-data that are ranked top in the current word segmentation sub-data ranking table and define them as keywords.
6. The data management method according to claim 1, characterized in that: Step S6, classifying all keywords according to the sentiment polarity and occurrence timestamp, and obtaining a sentiment polarity classification table based on a preset time period, including: Step S61, obtaining all keywords within the current preset time period and defining the keyword data set of the current preset time period , For the word segmentation data, is the number of all word segmentation sub-data; Step S62: defining a set of emotion polarity categories according to a plurality of preset types of emotion polarity ,in, For the category set The Emotional polarity, is the number of types of all sentiment polarities; Step S63, calculate the first The conditional probability of each word segmentation data under each sentiment polarity: (1); in, For the The conditional probability of keyword dataset under sentiment polarity; For the The marginal probability of a sentiment polarity; For the Emotional polarity The conditional probability of word segmentation data; Step S64, classifying each word segmentation sub-data into the sentiment polarity with the highest conditional probability; Step S65, integrating all classified word segmentation sub-data to obtain a sentiment polarity classification table for the current preset time period.
7. The data management method according to claim 1, characterized in that: Step S9, querying, updating and deleting data on the public opinion database of tin and indium materials through the SQL query statement, SQL update statement and SQL delete statement of the MySQL client, and then comprising: Step S100, generating a query record with a query timestamp and query content based on each data query; Step S200, generating an update record with an update timestamp and update content based on each data update; Step S300, generating a deletion record with a deletion timestamp and deletion content based on each data deletion; Step S400, sending all query records, all update records, and all deletion records to an external monitoring terminal.
8. A data management device for public opinion on tin and indium materials, the data management device being applied to the data management method according to any one of claims 1 to 7, characterized in that: The data management device comprises: The global environment variable file building module is used to build the global environment variable file of TiDB through the TiUP component manager; A TiDB cluster generation module, used to generate a TiDB cluster in the global environment variable file through the playground instruction of the TiUP component manager; A TiDB cluster connection module, used to connect to the TiDB cluster through a MySQL client signal; A tin-indium material public data crawling module, used to crawl a number of public data related to tin-indium materials on the public webpage through a data crawler script; The keyword extraction module is used to extract features from all public data and obtain several keywords; The sentiment polarity classification table acquisition module is used to classify all keywords according to the sentiment polarity and occurrence timestamp, and obtain a sentiment polarity classification table based on a preset time period; A blank database creation module, used to create a blank database with the same dimension as the sentiment polarity classification table in the TiDB cluster through the SQL creation statement of the MySQL client; A tin indium material public opinion database generation module is used to import all sentiment polarity classification tables into the blank database through the SQL import statement of the MySQL client to form a tin indium material public opinion database; The tin and indium material public opinion database management module is used to query, update and delete data in the tin and indium material public opinion database through the SQL query statement, SQL update statement and SQL delete statement of the MySQL client.
9. An electronic device, characterized in that: It comprises a processor and a memory coupled to the processor, wherein the memory stores program instructions executable by the processor; when the processor executes the program instructions stored in the memory, the data management method as described in any one of claims 1 to 7 is implemented.
10. A storage medium, characterized in that: The storage medium stores program instructions, and when the program instructions are executed by the processor, the data management method according to any one of claims 1 to 7 can be implemented.