Text analysis method and system based on vector database

The vector database is constructed through GloVe and BERT, combined with the SVM classifier and visualization platform, and the existing text analysis is solved, and the rapid multi-dimensional text analysis and intuitive report output are achieved.

CN120372003APending Publication Date: 2025-07-25ZHUO SHI TECH (HAINAN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410253241.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-06
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing text analysis technology relies on manual dominance, intelligent analysis is inefficient, and cannot effectively output multi-dimensional data features, and the user experience is unfriendly.

Method used

A vector database was constructed using GloVe and BERT Gemini model training, and text vector feature annotation and classification recognition were performed through SVM linear classifiers, and combined with the visual platform to output an analysis report.

Benefits of technology

It realizes fast and multi-dimensional text analysis, improves analysis efficiency, and users can intuitively obtain key text features and support large-scale high-capacity analysis tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372003A_ABST
    Figure CN120372003A_ABST
Patent Text Reader

Abstract

The invention provides a text analysis method and system based on a vector database, and relates to the technical field of intelligent big data processing. The method comprises the following steps: collecting and preprocessing historical big data of a text; the method comprises the following steps: respectively training and identifying historical big data of a text through double sub-models, namely GloVe and BERT, and fusing to construct a vector database; through text feature labeling, feature labeling on each dimension is carried out on each text vector in the vector database, and labeling features of each text vector on each dimension are obtained; using an SVM linear classifier to train and learn the labeling features of each text vector in each dimension to obtain a classification recognition model; and deploying the classification and recognition model on a visual platform for text recognition and analysis, and outputting a corresponding text analysis visual report. According to the method, the text can be quickly analyzed, and the analysis efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent big data processing, and particularly to a text analysis method and system based on a vector database. Background Art

[0002] Text analysis mainly involves collecting, screening, analyzing, and reporting relevant information. Text analysis is a complex task that requires professional skills and experience.

[0003] In current text analysis technologies, the traditional analysis method still mainly relies on manual work supplemented by intelligence. While text intelligent analysis technologies mainly use text analysis models based on deep learning algorithms. These algorithms are used to identify text data features, build models and apply them to text assisted analysis. However, in practice, due to insufficient text dimension annotation, it is unable to effectively output a large number of data vector features, the analysis tasks it can undertake are relatively low, it is not user-friendly, and the analysis efficiency is low. Summary of the Invention

[0004] To solve the above technical problems existing in the prior art, embodiments of the present invention provide a text analysis method and system based on a vector database. The technical solutions are as follows: On the one hand, a text analysis method based on a vector database is provided. The method includes the following steps: S1. Collect and preprocess text historical big data; S2. Respectively train and identify text historical big data through the twin models: GloVe and BERT, and fuse them to build a vector database; S3. Through text feature annotation, perform feature annotation on each dimension of each text vector in the vector database to obtain the annotation features of each text vector in each dimension; S4. Use an SVM linear classifier to train and learn the annotation features of each text vector in each dimension to obtain a classification and recognition model; S5. Deploy the classification and recognition model on a visualization platform for text recognition analysis and output a corresponding text analysis visualization report.

[0005] Optionally, S2. Respectively train and identify text historical big data through the twin models: GloVe and BERT, and fuse them to build a vector database, including: Pre-deploy the twin models: GloVe and BERT, that is, the GloVe neural network model and the BERT text vector model; The backend imports the text historical big data into the GloVe neural network model, and uses the GloVe neural network model to perform text vector statistics to obtain a corresponding text vector set M1; The backend imports the text historical big data into the BERT text vector model, uses the BERT text vector model to learn text vector features, identifies and extracts the corresponding text vectors, and obtains the corresponding text vector set M2; Fuse the text vectors in the text vector set M1 and the text vector set M2 and place them in the same database, and based on the clustering algorithm, cluster the text vectors with similarity features in the database to obtain several text vector subsets on the same dimension: {{m1}, {m2}, {m3}......}, where {m1}, {m2}, {m3}...... are text vector subsets on each dimension respectively; Adopt the method of vector average fusion to perform vector fusion and mean calculation on each text vector subset in {{m1}, {m2}, {m3}......}, obtain the text vectors on each dimension and store them in the database, so as to obtain a vector database.

[0006] Optionally, the step of obtaining the text vectors on each dimension and storing them in the database further includes: Assign storage ids to the text vectors on each dimension; Write the storage ids into the index retrieval data table preset in the backend; When new text big data appears, process the new text big data through the twin models GloVe and BERT to obtain new text vectors, and update the index retrieval data table in real time.

[0007] Optionally, in S3, through text feature annotation, perform feature annotation on each dimension of each text vector in the vector database to obtain the annotation features of each text vector on each dimension, including: Prepare text labels for text feature annotation in advance, and the text labels contain text features on each dimension; Import the text labels into the vector database, and automatically match the text labels on each dimension for each text vector in the vector database according to the dimension; After the matching is completed, mark and bind the text labels to the corresponding text vectors, so as to obtain the annotation features of each text vector on each dimension.

[0008] Optionally, in S4, use the SVM linear classifier to perform training and learning on the annotation features of each text vector on each dimension to obtain a classification and recognition model, including: Deploy the SVM linear classifier at the backend of the visualization platform; Send a data request to the vector database, requesting to import the text vectors in the vector database into the SVM linear classifier, where the storage id of the text vectors to be imported is included in the data request; Retrieve according to the storage id in the index retrieval data table, index the corresponding text vectors, and import them into the SVM linear classifier; Through the SVM linear classifier, train and learn the labeled features of the imported text vectors in each dimension to obtain the classification and recognition model.

[0009] Optionally, in S5, deploy the classification and recognition model to the visualization platform for text recognition analysis, and output the corresponding text analysis visualization report, including: Configure the parameters of the classification and recognition model. After the configuration is completed, deploy the classification and recognition model to the visualization platform for text recognition analysis; Through the front end of the visualization platform, access the text analysis data and transmit it to the back end of the visualization platform; The back end parses the text analysis data, calls the classification and recognition model, and indexes the text vectors and their labeled features in the corresponding dimension from the vector database according to the text analysis dimension; Write the text vectors and their labeled features indexed by the classification and recognition model into a preset visualization report to obtain the text analysis visualization report, and output it to the front end.

[0010] On the other hand, a system for implementing the text analysis method based on the vector database is provided, including the front end and the back end of the visualization platform, where, The front end is used to access the text analysis data and transmit it to the back end; The back end is used to index the text vectors and their labeled features in the corresponding dimension from the vector database through the deployed classification and recognition model according to the text analysis dimension; write the text vectors and their labeled features indexed by the classification and recognition model into a preset visualization report to obtain the text analysis visualization report, and output it to the front end; The front end is further used to visually display the text analysis visualization report; The front end is communicatively connected to the back end.

[0011] On the other hand, an electronic device is provided, and the electronic device includes: A processor; A memory, on which computer-readable instructions are stored, and when the computer-readable instructions are loaded and executed by the processor, the steps of the text analysis method as described above are implemented.

[0012] On the other hand, a computer-readable storage medium is provided, in which program codes are stored, and the program codes can be called by a processor to execute the steps of the above text analysis method.

[0013] The beneficial effects brought by the technical solution provided by the present invention at least include: The present invention constructs a vector database for identifying text vectors through training with the twin models: GloVe and BERT. Meanwhile, through text vector feature annotation in each dimension, a classification and recognition model for identifying text vector features is constructed. By combining the classification and recognition model with the vector database, the backend can quickly classify and recognize the large text data input from the frontend and index the corresponding text vectors and their annotation features in each dimension, and can quickly output the corresponding text vectors and annotation features in each dimension to the frontend, enabling the frontend user to quickly obtain effective multi-dimensional data vector features from the large text data, achieving the purpose of quickly analyzing the text and improving the analysis efficiency. Through the classification and recognition model, the text vectors can be updated, undertaking large-scale and high-capacity analysis tasks, making it more convenient for users.

[0014] At the same time, a visualization platform is adopted to process text analysis tasks, which can output a visualization analysis report, intuitively display the corresponding text vector features for users, and quickly obtain the key points of the text vectors. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0016] Figure 1 is a flowchart of a text analysis method based on a vector database provided by an embodiment of the present invention; Figure 2 is a schematic diagram of text recognition of the twin models provided by an embodiment of the present invention; Figure 3 is a schematic diagram of the component process of the vector database provided by an embodiment of the present invention; Figure 4 is an application schematic diagram of feature annotation provided by an embodiment of the present invention; Figure 5 is a schematic diagram of an application system provided by an embodiment of the present invention; Figure 6 is an application schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0018] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as an "example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of the word "example" is intended to present concepts in a specific manner. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.

[0019] In addition, to better illustrate the present invention, numerous specific details are given in the following specific implementation manners. Those skilled in the art should understand that the present invention can also be implemented without some specific details. In some instances, means well-known to those skilled in the art are not described in detail so as to highlight the gist of the present invention. Embodiment

[0020] As Figure 1 shown, on the one hand, the present invention provides a text analysis method based on a vector database, including the following steps: S1. Collect and preprocess big data of text history; S2. Respectively train to recognize big data of text history through the twin models: GloVe and BERT, and fuse and construct a vector database; S3. Through text feature annotation, perform feature annotation on each dimension of each text vector in the vector database to obtain the annotation features of each text vector in each dimension; S4. Use an SVM linear classifier to perform training and learning on the annotation features of each text vector in each dimension to obtain a classification and recognition model; S5. Deploy the classification and recognition model to a visualization platform for text recognition and analysis, and output a corresponding text analysis visualization report.

[0021] This solution aims to use GloVe and BERT for training to construct a model for identifying text vectors and building a vector database. At the same time, through the annotation of text vector features in each dimension, a classification and recognition model for identifying text vector features is constructed. By combining the classification and recognition model with the vector database, the backend can quickly classify and recognize the large text data input from the front end and index the corresponding text vectors and their annotation features in each dimension, and can quickly output the corresponding text vectors and annotation features in each dimension to the front end, enabling the front-end users to quickly obtain effective multi-dimensional data vector features from the large text data.

[0022] This solution is implemented on a visualization platform. Front-end users can submit large text data for real-time analysis to the backend and transmit it to the backend. The vector database and classification and recognition model deployed on the backend are used to analyze the large text data in terms of dimensions and output the corresponding text analysis report, and then feedback it to the front end. Users can view the visualized text analysis report on the front end, so as to quickly provide users with visualized text vector feature analysis, quickly obtain the corresponding text vector features from the large text data, and obtain the corresponding key text data.

[0023] As Figure 2 and Figure 3 shown, as an optional implementation of the present invention, optionally, S2. Through the twin models: GloVe and BERT, train and recognize the large text data of the history respectively, and fuse and build a vector database, including: Pre-deploy the twin models: GloVe and BERT, that is, the GloVe neural network model and the BERT text vector model; The backend imports the large text data of the history into the GloVe neural network model, and uses the GloVe neural network model to perform text vector statistics to obtain the corresponding text vector set M1; The backend imports the large text data of the history into the BERT text vector model, and uses the BERT text vector model to learn text vector features, identify and extract the corresponding text vectors to obtain the corresponding text vector set M2; Fuse the text vectors in the text vector set M1 and the text vector set M2 and place them in the same database, and based on the clustering algorithm, cluster the text vectors with similarity features in the database to obtain several subsets of text vectors in the same dimension: {{m1}, {m2}, {m3}......}, where {m1}, {m2}, {m3}...... are the subsets of text vectors in each dimension respectively; In the vector average fusion method, each subset of text vectors in {{m1}, {m2}, {m3}......} is subjected to vector fusion and mean calculation to obtain text vectors in each dimension and store them in the database, thereby obtaining a vector database.

[0024] The GloVe model, i.e., the Global Vectors model, believes that the statistics of word occurrences in the corpus (co-occurrence matrix) are important materials for unsupervised learning algorithms to learn word vector representations. The GloVe model uses the statistical information of the global (entire) corpus to generate the vector representation of this word.

[0025] In order to accurately output and identify text vector features, this solution also uses the BERT text vector model as an auxiliary. Its neural network layer identifies and classifies text word vectors from large text data, and together with the text vectors identified by the GloVe model, jointly constructs a text vector dataset for text vector identification and classification.

[0026] The specific steps are as follows: First, collect and preprocess large text data history.

[0027] 1. Data collection: Collect a large amount of text data, which will serve as the basis for training and constructing the text vector library.

[0028] 2. Data preprocessing: Clean and preprocess the collected data, including removing irrelevant content, unifying formats, word segmentation, removing stop words, etc., for subsequent processing.

[0029] 3. Establish a basic vector library: Use the preprocessed data to train the basic vector library. Common neural network models such as Word2Vec and GloVe can be used to train word vectors, or pre-trained models such as BERT and GPT can be used.

[0030] As Figure 3 shown, by using the twin models to respectively perform statistics on the large text data history for text vectors, corresponding text vector sets M1 and M2 can be obtained. By using the twin models to respectively extract and generate text vector sets M1 and M2 that are the same or similar, the text vectors identified and extracted by the two models can be integrated to obtain a more accurate and complete text vector dataset.

[0031] Therefore, the text vectors generated by the two models are placed in a database to obtain all the text vectors recognized and output by the two models and perform clustering. The text vectors with similarity features in the library are clustered by a clustering algorithm to obtain text vectors on the same dimension. For the clustering algorithm, this solution does not make a limitation. For the similarity features, a vector similarity calculation method can be adopted, such as cosine similarity. By calculating the similarity, the text vectors with similarity features are clustered, and the similar text vectors are used as subsets of text vectors on the same dimension. By the clustering algorithm and calculating the similarity, subsets of text vectors on different dimensions can be obtained. For example, the subset of text vectors on a certain dimension are all text vectors on a certain positive event. Through this positive event, the text vectors with the similarity features of this positive event are clustered to generate a subset of text vectors on the positive event dimension.

[0032] In the subset of text vectors, there are several text vectors with similarity features. The mean value of each vector is calculated by using the vector average fusion method, such as using the vector product quantization method for mean value calculation (vectorization), so as to obtain the text vectors on each dimension. That is to say, the mean value of several values in the subset of text vectors is calculated, but here the vector mean value calculation method is adopted to obtain the representative text vectors of each subset of text vectors.

[0033] Through the above processing, text vectors on different dimensions are obtained, and a vector database is constructed (stored in the library) and put into use.

[0034] As an optional implementation scheme of the present invention, optionally, obtaining the text vectors on each dimension and storing them in the library further includes: Allocating storage ids for the text vectors on each dimension; Writing the storage ids into a preset index retrieval data table at the back end; When new text big data appears, the new text big data is processed by the twin models GloVe and BERT to obtain new text vectors, and the index retrieval data table is updated in real time.

[0035] Specifically in storage, storage ids are allocated for the text vectors on each dimension, and the storage ids are written into a preset text vector id list, which is convenient for indexing the corresponding text vectors according to the ids later.

[0036] The text vector id list can be set according to the index retrieval method.

[0037] Subsequently, by retrieving the data table through the index, the text vector corresponding to the corresponding id can be found by traversing the index to retrieve the data table, which is convenient for outputting the corresponding text vector and improving the retrieval efficiency.

[0038] For the text vectors obtained from the new text data, they can be written into the index retrieval data table in the above manner. The index retrieval data table is stored in the backend database.

[0039] As Figure 4 shown, as an optional implementation of the present invention, optionally, S3. Through text feature annotation, perform feature annotation on each text vector in the vector database in each dimension to obtain the annotation features of each text vector in each dimension, including: Prepare in advance the text labels for text feature annotation, and the text labels include the text features in each dimension; Import the text labels into the vector database, and automatically match the text labels in each dimension for each text vector in the vector database according to the dimension; After the matching is completed, mark and bind the text labels to the corresponding text vectors, so as to obtain the annotation features of each text vector in each dimension.

[0040] Text annotation: Perform text annotation on the trained word vectors, and the annotation can be carried out manually or semi-automatically. The annotation labels can include sentiment polarity (positive, negative, neutral), sentiment intensity, etc.

[0041] The text labels in different dimensions are prepared by the administrator. It is convenient to identify and output the annotation features in the corresponding dimension after subsequent annotation. For example, output the corresponding annotation features for a certain text event, which is convenient for users to quickly obtain text keywords.

[0042] As Figure 5 shown, as an optional implementation of the present invention, optionally, S4. Use the SVM linear classifier to perform training and learning on the annotation features of each text vector in each dimension to obtain a classification and recognition model, including: Deploy the SVM linear classifier at the backend of the visualization platform; Send a data request to the vector database, requesting to import the text vectors in the vector database into the SVM linear classifier, where the data request includes the storage id of the text vectors to be imported; According to the storage id, retrieve in the index retrieval data table, index the corresponding text vector, and import it into the SVM linear classifier; Through the SVM linear classifier, perform training and learning on the annotation features of the imported text vectors in each dimension to obtain the classification and recognition model.

[0043] The SVM linear classifier is the SVM (Support Vector Machine). In this solution, the SVM (Support Vector Machine) is used to identify and learn the labeled features. The administrator can, according to the project requirements, index the text vectors in the corresponding dimension, input them into the SVM (Support Vector Machine) for identification and learning, and thus train the classifier directionally. The administrator can send a data request for a specific dimension to the database through the backend, index the text vectors in the corresponding dimension from the vector database, traverse and index them according to the id, and import them into the SVM (Support Vector Machine) for training and learning, so as to obtain a classification and recognition model. The classification and recognition model has the ability to identify and output the features of this dimension.

[0044] SVM (Support Vector Machine) is a commonly used machine learning algorithm, mainly used for classification and regression analysis. In classification problems, the SVM attempts to find a hyperplane that can maximize the separation of data points of different classes.

[0045] The linear classifier is a special form of the SVM, applicable to linearly separable data sets. In the linear classifier, the SVM attempts to find a hyperplane that can completely separate the data points of different classes. This hyperplane is defined by a normal vector and a threshold, where the normal vector determines the direction of the hyperplane and the threshold determines the position of the hyperplane.

[0046] For non-linear data sets, the SVM can map the data to a higher-dimensional space by using a kernel function, and then find a linearly separable hyperplane in this higher-dimensional space. Common kernel functions include linear kernel, polynomial kernel, and radial basis function (RBF).

[0047] The SVM linear classifier has a wide range of applications in many fields, such as text classification, image recognition, bioinformatics, and financial prediction, etc. It is a powerful and flexible machine learning algorithm, with the characteristics of being simple, intuitive, and easy to implement.

[0048] As an optional implementation of the present invention, optionally, in S5, the classification and recognition model is deployed on a visualization platform for text recognition analysis, and a corresponding text analysis visualization report is output, including: Configure the parameters of the classification and recognition model. After the configuration is completed, deploy the classification and recognition model on the visualization platform for text recognition analysis; Through the front end of the visualization platform, access the text analysis data and transmit it to the backend of the visualization platform; The backend parses the text analysis data, calls the classification and recognition model, and indexes the text vectors and their labeled features in the corresponding dimension from the vector database according to the text analysis dimension; Write the text vectors indexed by the classification and recognition model and their annotation features into a preset visualization report to obtain the text analysis visualization report and output it to the front end.

[0049] Deploy the model on the back end of the visualization platform for text recognition and analysis. After the front-end user inputs corresponding text big data to the back end through the front end, corresponding text vector analysis tasks can be established through each back end, and the classification and recognition model is scheduled to parse the current text analysis data, and the model is used to recognize the text analysis data. Through data parsing, the analysis dimensions required by the current front-end user can be obtained, and the text vectors and their annotation features corresponding to the dimensions can be indexed from the vector database according to the analysis dimensions.

[0050] The model can index the text vectors and annotation features corresponding to the text analysis dimensions by traversing the form according to the dimensions. Through the output text vectors and annotation features, quickly guide and extract the keyword information in the current text event for the front-end user.

[0051] In this solution, through the visualization report preset by the visualization platform, the output running vectors and annotation features are written into the preset visualization report, and the text analysis visualization report is output and sent to the front end. After being rendered by the front end, the visual text analysis report can be intuitively viewed to achieve visual display.

[0052] In this embodiment, the visualization platform is not limited.

[0053] Obviously, those skilled in the art should understand that to implement all or part of the processes in the above embodiments, it can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above control embodiments. Those skilled in the art can understand that to implement all or part of the processes in the above embodiments, it can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above control embodiments. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk (Hard Disk Drive, abbreviation: HDD) or a solid-state drive (Solid-State Drive, SSD), etc.; the storage medium can also include a combination of the above types of memories.

[0054] Embodiment 2 Based on the implementation principle of Embodiment 1, on the other hand, the present invention provides a system for implementing the text analysis method based on a vector database, including the front end and the back end of a visualization platform. Among them, The front end is used to access text analysis data and transmit it to the back end; The back end is used to index text vectors and their annotation features in corresponding dimensions from the vector database according to text analysis dimensions through the deployed classification and recognition model; write the text vectors and their annotation features indexed by the classification and recognition model into a preset visualization report to obtain the text analysis visualization report, and output it to the front end; The front end is further used to visually display the text analysis visualization report; The front end is communicatively connected to the back end.

[0055] For the above interaction control on the front end and the back end of the visualization platform, please specifically understand it in combination with the implementation principle of Embodiment 1, and it will not be elaborated in this embodiment.

[0056] Each module or each step of the above-mentioned present invention can be implemented by a general-purpose computing system. They can be concentrated on a single computing system or distributed on a network composed of multiple computing systems. Optionally, they can be implemented by program codes executable by the computing system. Thus, they can be stored in a storage system for execution by the computing system, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module for implementation. In this way, the present invention is not limited to any specific combination of hardware and software.

[0057] Embodiment 3 As Figure 6 shown, further, on the other hand, the present invention also provides an electronic device, including: A processor; A memory for storing processor-executable instructions; Wherein, the processor is configured to implement the text analysis method based on a vector database when executing the executable instructions.

[0058] The electronic device in the embodiment of the present invention includes a processor and a memory for storing processor-executable instructions. Among them, the processor is configured to implement the text analysis method based on a vector database described in any one of the foregoing when executing the executable instructions.

[0059] Here, it should be noted that the number of processors can be one or more. At the same time, in the electronic device according to the embodiment of the present invention, an input system and an output system may also be included. Among them, the processor, the memory, the input system and the output system may be connected through a bus or in other ways, which is not specifically limited herein.

[0060] As a computer-readable storage medium, the memory can be used to store software programs, computer-executable programs and various modules, such as: the programs or modules corresponding to a text analysis method based on a vector database according to the embodiment of the present invention. The processor executes various functional applications and data processing of the electronic device by running the software programs or modules stored in the memory.

[0061] The input system can be used to receive input numbers or signals. Among them, the signal can be a key signal related to the user settings and function control of the device / terminal / server. The output system may include a display device such as a display screen.

[0062] In an exemplary embodiment, the present invention also provides a computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is loaded and executed by the processor to implement the steps of the text analysis method based on a vector database as described above. For example, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0063] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or terminal device including the element.

[0064] When "an embodiment", "the embodiment", "an exemplary embodiment", "some embodiments" etc. are mentioned in the specification, it indicates that the described embodiment may include specific features, structures or characteristics, but not necessarily every embodiment includes the specific feature, structure or characteristic. In addition, when combining an embodiment to describe a specific feature, structure or characteristic, implementing such feature, structure or characteristic in combination with other embodiments (whether explicitly described or not) should be within the knowledge of those skilled in the relevant art.

[0065] It should be understood that the term "and / or" in this text is merely a relational description of associated objects, indicating that three relationships can exist. For example, A and / or B can represent three cases: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. Additionally, the character " / " in this text generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship. Specific understanding can be referred to the context before and after.

[0066] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single item(s) or plural item(s). For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0067] It should be understood that in various embodiments of the present invention, the magnitude of the serial numbers of the above processes does not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0068] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.

[0069] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0070] In addition, in each embodiment of the present invention, the functional units can be integrated in one processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0071] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0072] The present invention covers any alternatives, modifications, equivalent methods, and solutions made within the spirit and scope of the present invention. For the public to have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention. However, those skilled in the art can fully understand the present invention without the description of these details. In addition, well-known methods, processes, procedures, components, and circuits, etc., are not described in detail to avoid unnecessary confusion to the essence of the present invention.

[0073] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc., made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A text analysis method based on a vector database, characterized in that, It includes the following steps: S1. Collect and preprocess the large historical text data; S2. Respectively train to recognize the large historical text data through the twin models: GloVe and BERT, and fuse them to build a vector database; S3. Through text feature annotation, perform feature annotation on each text vector in the vector database in each dimension to obtain the annotation features of each text vector in each dimension; S4. Use the SVM linear classifier to train and learn the annotation features of each text vector in each dimension to obtain a classification recognition model; S5. Deploy the classification recognition model to a visualization platform for text recognition analysis and output the corresponding text analysis visualization report.

2. The text analysis method based on a vector database according to claim 1, wherein S2. Respectively train to recognize the large historical text data through the twin models: GloVe and BERT, and fuse them to build a vector database, including: Pre-deploy the twin models: GloVe and BERT, namely the GloVe neural network model and the BERT text vector model; The backend imports the large historical text data into the GloVe neural network model, and uses the GloVe neural network model to perform text vector statistics to obtain the corresponding text vector set M1; The backend imports the large historical text data into the BERT text vector model, and uses the BERT text vector model to learn text vector features, recognize and extract the corresponding text vectors to obtain the corresponding text vector set M2; Fuse the text vectors in the text vector set M1 and the text vector set M2 into the same database, and based on the clustering algorithm, cluster the text vectors with similarity features in the database to obtain several subsets of text vectors in the same dimension: {{m1}, {m2}, {m3}......}, where {m1}, {m2}, {m3}...... are respectively the subsets of text vectors in each dimension; Adopt the vector average fusion method to perform vector fusion and mean calculation on each subset of text vectors in {{m1}, {m2}, {m3}......} to obtain the text vectors in each dimension and store them in the database, thereby obtaining the vector database.

3. The text analysis method based on a vector database according to claim 2, wherein, The step of obtaining the text vectors in each dimension and storing them in the database further includes: Assign storage ids to the text vectors in each dimension; Write the storage ids into the index retrieval data table preset in the backend; When new text big data appears, process the new text big data through the twin models GloVe and BERT to obtain new text vectors and update the index retrieval data table in real time.

4. A text analysis method based on a vector database according to claim 1, characterized in that, S3. Through text feature annotation, perform feature annotation on each text vector in the vector database in each dimension to obtain the annotation features of each text vector in each dimension, including: Prepare in advance the text labels for text feature annotation, and the text labels contain the text features in each dimension; Import the text labels into the vector database and automatically match the text labels in each dimension for each text vector in the vector database according to the dimension; After matching, the text tags are marked and bound to the corresponding text vectors, thereby obtaining the labeled features of each text vector in each dimension.

5. A text analysis method based on a vector database according to claim 1, characterized in that, S4. Using an SVM linear classifier, training and learning the labeled features of each text vector in each dimension to obtain a classification and recognition model, including: Deploying an SVM linear classifier at the backend of the visualization platform; Sending a data request to the vector database, requesting to import the text vectors in the vector database into the SVM linear classifier, where the data request includes the storage id of the text vectors to be imported; According to the storage id, retrieving in the index retrieval data table to index the corresponding text vectors and importing them into the SVM linear classifier; Through the SVM linear classifier, training and learning the labeled features of the imported text vectors in each dimension to obtain the classification and recognition model.

6. A text analysis method based on a vector database according to claim 5, characterized in that, S5. Deploying the classification and recognition model on the visualization platform for text recognition analysis and outputting a corresponding text analysis visualization report, including: Performing parameter configuration on the classification and recognition model, and after the configuration is completed, deploying the classification and recognition model on the visualization platform for text recognition analysis; Through the front end of the visualization platform, accessing text analysis data and transmitting it to the backend of the visualization platform; The backend parses the text analysis data and calls the classification and recognition model, and indexes the text vectors and their labeled features in the corresponding dimension from the vector database according to the text analysis dimension; Writing the text vectors and their labeled features indexed by the classification and recognition model into a preset visualization report to obtain the text analysis visualization report and outputting it to the front end.

7. A system for implementing the text analysis method based on a vector database according to any one of claims 1-6, characterized in that, It includes the front end and the backend of the visualization platform, where The front end is used to access text analysis data and transmit it to the backend; The backend is used to index the text vectors and their labeled features in the corresponding dimension from the vector database according to the text analysis dimension through the deployed classification and recognition model; writing the text vectors and their labeled features indexed by the classification and recognition model into a preset visualization report to obtain the text analysis visualization report and outputting it to the front end; The front end is further used to visually display the text analysis visualization report; The front end and the backend are communicatively connected.

8. An electronic device, characterized in that, The electronic device includes: A processor; A memory, on which computer-readable instructions are stored, and when the computer-readable instructions are loaded and executed by the processor, the method according to any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium, characterized in that, Program code is stored in the computer-readable storage medium, and the program code can be called by the processor to execute the method according to any one of claims 1 to 6.