Teaching field data acquisition method based on data medium table and related device
By receiving the user's data retrieval needs in the data, generating embedded word vectors and calling large language models, the redundancy and accuracy problems of data collection in colleges and universities are solved, and efficient data retrieval and collection are achieved.
Patent Information
- Application Number
- CN202510632108.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-06-24
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
During the digital transformation of colleges and universities, the teaching field data collected in Taichung is prone to redundancy and impurities when collected manually, and it is difficult to ensure data accuracy and collection efficiency.
By receiving user data in the data searching for demand text, extracting demand semantic information, generating embedded word vectors, and performing similarity calculations with embedded word vectors in the index database, selecting corresponding embedded word vectors, constructing prompt words, and calling large language models to accurately retrieve teaching field data.
It realizes accurate retrieval of required teaching data in Taichung in the data, improves the efficiency of data collection, and reduces the redundancy and errors of manual processing.
Smart Images

Figure CN120196726A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a method for collecting teaching field data based on a data middle platform and related devices. Background Art
[0002] When colleges and universities are undergoing digital transformation, all data is stored in the data middle platform, and colleges and universities carry out a large number of tasks such as data collection, statistics, verification, and analysis every year; therefore, various data will be converged and stored in the data middle platform. When relevant personnel need to query data in the corresponding teaching field, collecting data manually may result in a large amount of redundant or impure data, which requires a large amount of manpower and material resources for further data screening and processing, and cannot guarantee the accuracy of the collected data and the efficiency of data collection. Summary of the Invention
[0003] An object of the present invention is to overcome the deficiencies of the prior art. The present invention provides a method for collecting teaching field data based on a data middle platform and related devices, which can accurately retrieve the required teaching field data in the data middle platform, thereby improving the efficiency of data collection.
[0004] To solve the above technical problems, an embodiment of the present invention provides a method for collecting teaching field data based on a data middle platform, which is applied to data collection in the teaching field. The method includes: The data middle platform receives a data retrieval requirement text uploaded by a user based on a retrieval terminal, and extracts the retrieval requirement semantic information in the data retrieval requirement text; Perform semantic word segmentation processing on the retrieval requirement semantic information, and generate a retrieval requirement embedding word vector based on the word segmentation processing result of the retrieval requirement semantic information; Perform vector similarity calculation processing on the retrieval requirement embedding word vector and the embedding word vectors stored in the index database, and obtain the vector similarity calculation result of the retrieval requirement embedding word vector. The embedding word vectors stored in the index database are constructed using the text descriptions stored in the teaching domain data warehouse in the data middle platform; Select the corresponding retrieval embedding word vector based on the vector similarity calculation result; Construct a prompt word based on the retrieval embedding word vector, and call a large language model based on the prompt word to return the teaching field data information that needs to be collected in the data middle platform.
[0005] Optionally, the data middle platform receiving the data retrieval requirement text uploaded by the user based on the retrieval terminal includes: After the data middle platform establishes a communication connection with the retrieval terminal, it receives the retrieval permission authentication information uploaded by the user based on the retrieval terminal; The data middle platform authenticates the retrieval permission authentication information and assigns corresponding retrieval permissions to the user on the retrieval terminal based on the authentication result; The data middle platform receives the data retrieval requirement text uploaded by the user on the retrieval terminal based on the assigned corresponding retrieval permissions.
[0006] Optionally, the retrieved requirement semantic information extracted from the data retrieval requirement text includes: Call a corresponding semantic extraction model in the data middle platform, input the data retrieval requirement text into the semantic extraction model, and extract the retrieved requirement semantic information in the data retrieval requirement text based on the semantic extraction model; Wherein the semantic extraction model is a model obtained by training and converging in a semantic expression extraction network by using historical data retrieval requirement texts and the expert-labeled semantic content corresponding to the historical data retrieval requirement texts.
[0007] Optionally, the semantic word segmentation process is performed on the retrieved requirement semantic information, and a retrieved requirement embedded word vector is generated based on the word segmentation result of the retrieved requirement semantic information, including: Perform semantic word segmentation on the retrieved requirement semantic information in the way of semantic word segmentation to form a word segmentation result corresponding to the retrieved requirement semantic information. The word segmentation result is to segment the retrieved requirement semantic information into a series of token word segments. The token word segment is the basic word segmentation unit of the retrieved requirement semantic information and is a single character or phrase; Based on a large language model, convert each word segment in the word segmentation result into a vector in an embedding matrix, and perform an accumulation process on the vectors to form a retrieved requirement embedded word vector. The embedding matrix is a vector library, and each vector in the vector library corresponds to a specific token word segment.
[0008] Optionally, the vector similarity calculation process is performed on the retrieved requirement embedded word vector and the embedded word vectors stored in the index database to obtain the vector similarity calculation result of the retrieved requirement embedded word vector, including: Use the cosine similarity algorithm to perform vector similarity calculation on the retrieved requirement embedded word vector and the embedded word vectors stored in the index database to obtain the first similarity calculation result of the retrieved requirement embedded word vector; Use the Euclidean distance similarity algorithm to perform vector similarity calculation on the retrieved requirement embedded word vector and the embedded word vectors stored in the index database to obtain the second similarity calculation result of the retrieved requirement embedded word vector; Perform a linear weighting process on the first similarity calculation result and the second similarity calculation result to obtain the vector similarity calculation result of the retrieval requirement embedded word vector.
[0009] Optionally, the selecting the corresponding retrieval embedded word vector based on the vector similarity calculation result includes: Perform a sorting process on the vector similarity calculation result, and select the corresponding retrieval embedded word vector based on the vector similarity sorting result.
[0010] Optionally, the constructing a prompt word based on the retrieval embedded word vector and calling a large language model based on the prompt word to return the teaching domain data information to be collected in the data center includes: Input the retrieval embedded word vector into a prompt word construction model, and generate the prompt word based on the prompt word construction model. The prompt word construction model is a pre-trained language model, and during the training process, the input training data is designed, experimented with, and optimized to guide the model to generate targeted prompt words; Call the large language model based on the prompt word, and input the prompt word and the retrieval embedded word vector into the large language model to perform a retrieval process on the teaching domain data in the data center to obtain the teaching domain data information to be collected.
[0011] In addition, an embodiment of the present invention further provides a teaching domain data collection device based on a data center, which is applied to the data collection in the teaching domain. The device includes: A receiving module: configured to receive, by the data center, a data retrieval requirement text uploaded by a user based on a retrieval terminal, and extract the retrieval requirement semantic information in the data retrieval requirement text; A generating module: configured to perform semantic word segmentation processing on the retrieval requirement semantic information, and generate a retrieval requirement embedded word vector based on the word segmentation processing result of the retrieval requirement semantic information; A similarity calculation module: configured to perform a vector similarity calculation process on the retrieval requirement embedded word vector and the embedded word vector stored in an index database to obtain the vector similarity calculation result of the retrieval requirement embedded word vector. The embedded word vector stored in the index database is constructed using the text descriptions stored in the teaching domain data warehouse in the data center; A selecting module: configured to select the corresponding retrieval embedded word vector based on the vector similarity calculation result; A data collection module: configured to construct a prompt word based on the retrieval embedded word vector, and call a large language model based on the prompt word to return the teaching domain data information to be collected in the data center.
[0012] In addition, an embodiment of the present invention further provides an electronic device, including a processor and a memory. The processor runs a computer program or code stored in the memory to implement the teaching field data acquisition method described in any one of the above.
[0013] In addition, an embodiment of the present invention further provides a computer-readable storage medium for storing a computer program or code. When the computer program or code is executed by a processor, the teaching field data acquisition method described in any one of the above is implemented.
[0014] In an embodiment of the present invention, by receiving a data retrieval requirement text, extracting the corresponding retrieval requirement semantic information in the data retrieval requirement text, constructing a retrieval requirement embedding word vector through the retrieval requirement semantic information; selecting a corresponding retrieval embedding sub-vector through the similarity with the embedding word vector stored in the index database, finally constructing a prompt word, and realizing the acquisition of corresponding teaching field data information in the data middle platform through a large language model; by utilizing semantic embedding and search technology in the large model, it is possible to accurately retrieve the required teaching field data in the data middle platform, thereby improving the efficiency of data acquisition. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1 is a schematic flowchart of a teaching field data acquisition method based on a data middle platform in an embodiment of the present invention; Figure 2 is a schematic flowchart of a teaching field data acquisition method based on a data middle platform in another embodiment of the present invention; Figure 3 is a schematic structural composition diagram of a teaching field data acquisition device based on a data middle platform in an embodiment of the present invention; Figure 4 is a schematic structural composition diagram of an electronic device in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0018] Embodiment 1, please refer to Figure 1 , Figure 1 which is a schematic flowchart of a data collection method for the teaching field based on a data middle platform in the embodiments of the present invention.
[0019] As Figure 1 shown, a data collection method for the teaching field based on a data middle platform is applied to data collection in the teaching field. The method includes: S101: The data middle platform receives a data retrieval requirement text uploaded by a user based on a retrieval terminal, and extracts retrieval requirement semantic information in the data retrieval requirement text; In the specific implementation process of the present invention, the data middle platform receiving the data retrieval requirement text uploaded by the user based on the retrieval terminal includes: after the data middle platform establishes a communication connection with the retrieval terminal, receiving the retrieval permission authentication information uploaded by the user based on the retrieval terminal; the data middle platform performs authentication processing on the retrieval permission authentication information, and assigns corresponding retrieval permissions to the user on the retrieval terminal based on the authentication result; the data middle platform receives the data retrieval requirement text uploaded by the user on the retrieval terminal based on the assigned corresponding retrieval permissions.
[0020] Further, the extracting the retrieval requirement semantic information in the data retrieval requirement text includes: calling a corresponding semantic extraction model in the data middle platform, inputting the data retrieval requirement text into the semantic extraction model, and extracting the retrieval requirement semantic information in the data retrieval requirement text based on the semantic extraction model; wherein the semantic extraction model is a model obtained by training and converging in a semantic expression extraction network by inputting historical data retrieval requirement texts and expert-labeled semantic contents corresponding to the historical data retrieval requirement texts.
[0021] Specifically, in the data middle platform, in order to ensure data security, for any operation that needs to perform data retrieval and collection on the data middle platform, corresponding permission authentication is required, and ultimately, data retrieval needs to be performed according to the assigned corresponding permissions, so as to collect teaching field data within the scope of the corresponding permissions, thereby ensuring data security.
[0022] Therefore, after establishing a communication connection between the data middle platform and the corresponding retrieval terminal, where the retrieval terminal is an intelligent terminal installed with a corresponding port or APP; the data middle platform will receive the retrieval permission authentication information uploaded by relevant users through the retrieval terminal; and authenticate the received retrieval permission authentication information, and after obtaining the authentication result, allocate the corresponding retrieval permission to the user on the retrieval terminal according to the authentication result; after the user obtains the corresponding retrieval permission on the retrieval terminal, the retrieval terminal will use the retrieval permission to receive the data retrieval requirement text within the retrieval permission range of the user, and upload the data retrieval requirement text to the data middle platform, so that the data middle platform can receive the data retrieval requirement text uploaded by the user.
[0023] After receiving the data retrieval requirement text, it is necessary to extract the retrieval requirement semantic information in the data retrieval requirement text. At this time, it is necessary to use a semantic extraction model to achieve this, that is, call the corresponding semantic extraction model in the data middle platform, and input the data retrieval requirement text into the semantic extraction model, and perform semantic extraction operations in the semantic extraction model, so as to extract the retrieval requirement semantic information in the data retrieval requirement text; where the semantic extraction model is a model formed by inputting historical data retrieval requirement texts and the expert-labeled semantic content corresponding to the historical data retrieval requirement texts into a semantic expression extraction network for training and converging after training.
[0024] S102: Perform semantic word segmentation processing on the retrieval requirement semantic information, and generate a retrieval requirement embedded word vector based on the word segmentation processing result of the retrieval requirement semantic information; In the specific implementation process of the present invention, the performing semantic word segmentation processing on the retrieval requirement semantic information and generating a retrieval requirement embedded word vector based on the word segmentation processing result of the retrieval requirement semantic information includes: performing semantic word segmentation processing on the retrieval requirement semantic information in the way of semantic word segmentation to form a word segmentation processing result corresponding to the retrieval requirement semantic information, and the word segmentation processing result is to segment the retrieval requirement semantic information into a series of marked word segments, and the marked word segments are the basic word segmentation units of the retrieval requirement semantic information, which are single characters or phrases; converting each word segment in the word segmentation processing result into a vector in an embedding matrix based on a large language model, and performing an accumulation process on the vectors to form a retrieval requirement embedded word vector, and the embedding matrix is a vector library, and each vector in the vector library corresponds to a specific marked word segment.
[0025] Specifically, it is necessary to perform semantic word segmentation on the semantic information of the retrieval requirement, and then construct a retrieval requirement embedding word vector through semantic word segmentation; when performing semantic word segmentation, it is necessary to perform the word segmentation operation in a semantic manner, that is, perform semantic word segmentation on the semantic information of the retrieval requirement according to semantics, so as to form a word segmentation result corresponding to the semantic information of the retrieval requirement, where the word segmentation result is to segment the semantic information of the retrieval requirement into a series of token word segments, and the token word segment is the basic word segmentation unit of the semantic information of the retrieval requirement, which is a single character or phrase.
[0026] After obtaining the word segmentation result, it is necessary to convert each word segment in the word segmentation result into a vector in the embedding matrix through a large language model (LLM), and then perform an accumulation process on the vectors to form a retrieval requirement embedding word vector, where the embedding matrix is a vector library, and each vector in the vector library corresponds to a specific token word segment.
[0027] There is a semantic embedding module in the large language model, that is, each word segment in the input word segmentation result is converted into a vector in the embedding matrix through the semantic embedding module in the large language model; and when performing the accumulation process on the vectors, the accumulation process is also performed in the large language model, mainly integrating into a single vector, which comprehensively represents the information input by the entire semantic text information of the retrieval requirement; this comprehensive vector is called a semantic embedding vector, that is, the retrieval requirement embedding word vector in this application; the retrieval requirement embedding word vector can capture the overall meaning and context information of the semantic text information of the retrieval requirement.
[0028] S103: Perform a vector similarity calculation process on the retrieval requirement embedding word vector and the embedding word vectors stored in the index database to obtain a vector similarity calculation result of the retrieval requirement embedding word vector, where the embedding word vectors stored in the index database are constructed using the text descriptions stored in the teaching domain data warehouse in the data middle platform; In the specific implementation process of the present invention, the performing a vector similarity calculation process on the retrieval requirement embedding word vector and the embedding word vectors stored in the index database to obtain a vector similarity calculation result of the retrieval requirement embedding word vector includes: performing a vector similarity calculation process on the retrieval requirement embedding word vector and the embedding word vectors stored in the index database using the cosine similarity algorithm to obtain a first similarity calculation result of the retrieval requirement embedding word vector; performing a vector similarity calculation process on the retrieval requirement embedding word vector and the embedding word vectors stored in the index database using the Euclidean distance similarity algorithm to obtain a second similarity calculation result of the retrieval requirement embedding word vector; performing a linear weighting process on the first similarity calculation result and the second similarity calculation result to obtain a vector similarity calculation result of the retrieval requirement embedding word vector.
[0029] Specifically, the data middle platform needs to compare the retrieval requirements formed by the input with the embedded word vectors stored in the embedded word vector index database. This can be achieved using similarity. In this application, the similarity of comparison vectors with different metrics can be adopted, and then comprehensive weighting can be performed to more accurately obtain the corresponding similarity. In this application, the cosine similarity algorithm and the Euclidean distance similarity algorithm can be used. In this embodiment, the higher the similarity or the smaller the distance, the stronger the semantic relationship between the vectors.
[0030] That is, the cosine similarity algorithm is used to perform vector similarity calculation processing on the retrieval requirement embedded word vector and the embedded word vector stored in the index database, so as to obtain the first similarity calculation result of the retrieval requirement embedded word vector; then the Euclidean distance similarity algorithm is used to perform vector similarity calculation processing on the retrieval requirement embedded word vector and the embedded word vector stored in the index database, so as to obtain the second similarity calculation result of the retrieval requirement embedded word vector; finally, linear weighting processing is performed on the first similarity calculation result and the second similarity calculation result to obtain the vector similarity calculation result of the retrieval requirement embedded word vector.
[0031] In this embodiment, regarding the index database, embedded word vectors are stored in the index database, where the embedded word vectors are constructed using the text descriptions stored in the teaching domain data warehouse in the data middle platform; that is, the embedded word vectors are generated using the text descriptions corresponding to the data such as application tables, summary tables, and detail tables in the teaching domain data warehouse.
[0032] S104: Select the corresponding retrieval embedded word vector based on the vector similarity calculation result; In the specific implementation process of the present invention, the selecting the corresponding retrieval embedded word vector based on the vector similarity calculation result includes: sorting the vector similarity calculation result, and selecting the corresponding retrieval embedded word vector based on the vector similarity sorting result.
[0033] Specifically, first, the vector similarity calculation result is sorted, and then the embedded word vectors corresponding to the top several similarities in the vector similarity sorting result are selected as the corresponding retrieval embedded word vectors.
[0034] S105: Construct a prompt word based on the retrieval embedded word vector, and call a large language model based on the prompt word to return the teaching domain data information that needs to be collected in the data middle platform.
[0035] In the specific implementation process of the present invention, constructing a prompt word based on the retrieved embedded word vector, and calling a large language model based on the prompt word to return the teaching domain data information to be collected in the data platform, including: inputting the retrieved embedded word vector into a prompt word construction model, and generating the prompt word based on the prompt word construction model, where the prompt word construction model is a pre-trained language model, and during the training process, the input training data is designed, experimented with, and optimized to guide the model to generate targeted prompt words; calling the large language model based on the prompt word, and inputting the prompt word and the retrieved embedded word vector into the large language model to perform retrieval processing on the teaching domain data in the data platform, so as to obtain the teaching domain data information to be collected.
[0036] Specifically, it is first necessary to construct a prompt word through the retrieved embedded word vector, and finally perform retrieval processing on the teaching domain data in the data platform through the prompt word and the retrieved embedded word vector, so as to obtain the teaching domain data information to be collected.
[0037] Therefore, it is necessary to input the retrieved embedded word vector into the prompt word construction model, and generate a prompt word through the prompt word construction model, where the prompt word construction model is a pre-trained language model, and during the training process, the input training data is designed, experimented with, and optimized to guide the model to generate targeted prompt words; after obtaining the prompt word, call the corresponding large language model through the prompt word, input the prompt word and the retrieved embedded word vector into the large language model, and perform retrieval processing on the teaching domain data in the data platform in the large language model, so as to obtain the teaching domain data information to be collected.
[0038] That is, the prompt word construction model is a pre-trained speech model, and a technology that guides the model to generate high-quality, accurate, and targeted outputs by designing, experimenting with, and optimizing the input prompt words; essentially, prompt engineering is also a way of human-computer interaction, and the prompt word is the input (instruction) provided to the LLM (large language model). The large model outputs content related to the instruction according to the instruction and its pre-trained "knowledge". The quality of the output result of the large model is related to the input instruction.
[0039] Based on the reference example of the prompt word of Text2Sql, the prompt word generally consists of the following elements: Role: Define a role for the large model that matches the target task; it can clearly define its role in one sentence (such as "You are a big data engineer"), so as to effectively narrow the problem domain, reduce ambiguity, and guide the large model to deepen from "general" to "professional field".
[0040] Constraint: Describe the specific task in detail.
[0041] Context: Provide other background information related to the task (such as historical conversations, scenarios, etc.).
[0042] Example: Examples are important and are an important reference when the large model generates output, which is very helpful for the output result.
[0043] Input: The input information of the task, which has a clear "input" identifier in the prompt.
[0044] Output: The format description of the output, which returns the result in JSON format, etc.
[0045] According to the Text2Sql solution based on the large model, it is necessary to design prompts for the undergraduate teaching domain.
[0046] In the embodiment of the present invention, by receiving the data retrieval requirement text, extracting the corresponding retrieval requirement semantic information in the data retrieval requirement text, and then constructing a retrieval requirement embedding word vector through the retrieval requirement semantic information; and selecting the corresponding retrieval embedding sub-vector through the similarity with the embedding word vector stored in the index database, and finally constructing a prompt, and realizing the acquisition of the corresponding teaching domain data information in the data middle platform through the large language model; by using the semantic embedding and search technology in the large model, it is possible to accurately retrieve the required teaching domain data in the data middle platform, thereby improving the efficiency of data acquisition.
[0047] Embodiment 2, please refer to Figure 2 , Figure 2 is a schematic flowchart of a method for collecting teaching domain data based on a data middle platform in another embodiment of the present invention.
[0048] As Figure 2 shown, a method for collecting teaching domain data based on a data middle platform, which is applied to the data collection in the teaching domain, the method includes: S201: The data middle platform receives the data retrieval requirement text uploaded by the user based on the retrieval terminal, and extracts the retrieval requirement semantic information in the data retrieval requirement text; S202: Perform semantic word segmentation processing on the retrieval requirement semantic information in the way of semantic word segmentation, form a word segmentation processing result corresponding to the retrieval requirement semantic information, the word segmentation processing result is to segment the retrieval requirement semantic information into a series of marked word segments, and the marked word segment is the basic word segmentation unit of the retrieval requirement semantic information, which is a single character or a phrase; S203: Based on the large language model, convert each word segment in the word segmentation processing result into a vector in the embedding matrix, and perform an accumulation process on the vectors to form a retrieval requirement embedding word vector, the embedding matrix is a vector library, and each vector in the vector library corresponds to a specific marked word segment; S204: Use the cosine similarity algorithm to perform vector similarity calculation processing on the retrieved demand embedding vector and the embedding vectors stored in the index database, and obtain the first similarity calculation result of the retrieved demand embedding vector; S205: Use the Euclidean distance similarity algorithm to perform vector similarity calculation processing on the retrieved demand embedding vector and the embedding vectors stored in the index database, and obtain the second similarity calculation result of the retrieved demand embedding vector; S206: Perform linear weighting processing on the first similarity calculation result and the second similarity calculation result to obtain the vector similarity calculation result of the retrieved demand embedding vector; S207: Select the corresponding retrieved embedding vector based on the vector similarity calculation result; S208: Construct a prompt word based on the retrieved embedding vector, and call a large language model based on the prompt word to return the teaching domain data information to be collected in the data middle platform.
[0049] For the specific implementation of the second embodiment, reference may be made to the first embodiment, which will not be elaborated here.
[0050] Embodiment 3, please refer to Figure 3 , a teaching domain data collection device based on a data middle platform, which is applied to the data collection in the teaching domain. The device includes: Receiving module 301: used to receive, by the data middle platform, the data retrieval demand text uploaded by the user based on the retrieval terminal, and extract the retrieval demand semantic information in the data retrieval demand text; In the specific implementation process of the present invention, the data middle platform receiving the data retrieval demand text uploaded by the user based on the retrieval terminal includes: after the data middle platform establishes a communication connection with the retrieval terminal, receiving the retrieval permission authentication information uploaded by the user based on the retrieval terminal; the data middle platform performs authentication processing on the retrieval permission authentication information, and assigns corresponding retrieval permissions to the user on the retrieval terminal based on the authentication result; the data middle platform receives the data retrieval demand text uploaded by the user on the retrieval terminal based on the assigned corresponding retrieval permissions.
[0051] Further, the extracting the retrieval demand semantic information in the data retrieval demand text includes: calling a corresponding semantic extraction model in the data middle platform, inputting the data retrieval demand text into the semantic extraction model, and extracting the retrieval demand semantic information in the data retrieval demand text based on the semantic extraction model; wherein the semantic extraction model is a model obtained by training and converging in a semantic expression extraction network by inputting historical data retrieval demand texts and the expert-labeled semantic content corresponding to the historical data retrieval demand texts.
[0052] Specifically, in the data center, to ensure data security, for any operation that needs to retrieve and collect data from the data center, corresponding permission authentication is required, and ultimately, data retrieval is performed according to the allocated corresponding permissions, so as to collect teaching field data within the corresponding permission range, thereby ensuring data security.
[0053] Therefore, after establishing a communication connection between the data center and the corresponding retrieval terminal, where the retrieval terminal is an intelligent terminal installed with corresponding ports or APPs; the data center will receive the retrieval permission authentication information uploaded by relevant users through the retrieval terminal; and perform authentication processing on the received retrieval permission authentication information, and after obtaining the authentication result, allocate corresponding retrieval permissions for the user on the retrieval terminal according to the authentication result; after the user obtains the corresponding retrieval permission on the retrieval terminal, the retrieval terminal will use the retrieval permission to receive the data retrieval requirement text within the user's retrieval permission range and upload the data retrieval requirement text to the data center, so that the data center can receive the data retrieval requirement text uploaded by the user.
[0054] After receiving the data retrieval requirement text, it is necessary to extract the retrieval requirement semantic information in the data retrieval requirement text. At this time, it is necessary to use a semantic extraction model to achieve this, that is, call the corresponding semantic extraction model in the data center, and input the data retrieval requirement text into the semantic extraction model, and perform semantic extraction operations in the semantic extraction model, so as to extract the retrieval requirement semantic information in the data retrieval requirement text; where the semantic extraction model is a model formed by inputting historical data retrieval requirement texts and the corresponding expert-labeled semantic content of the historical data retrieval requirement texts into a semantic expression extraction network for training and converging after training.
[0055] Generation module 302: used to perform semantic word segmentation processing on the retrieval requirement semantic information, and generate a retrieval requirement embedded word vector based on the word segmentation processing result of the retrieval requirement semantic information; In the specific implementation process of the present invention, the semantic tokenization processing of the semantic information of the retrieval requirement and the generation of the retrieval requirement embedding vector based on the tokenization result of the semantic information of the retrieval requirement include: performing semantic tokenization processing on the semantic information of the retrieval requirement in the way of semantic tokenization to form a tokenization result corresponding to the semantic information of the retrieval requirement, where the tokenization result is to segment the semantic information of the retrieval requirement into a series of tokenized words, and the tokenized words are the basic tokenization units of the semantic information of the retrieval requirement, which are single characters or phrases; based on a large language model, converting each tokenized word in the tokenization result into a vector in an embedding matrix, and performing an accumulation process on the vectors to form a retrieval requirement embedding vector, where the embedding matrix is a vector library, and each vector in the vector library corresponds to a specific tokenized word.
[0056] Specifically, it is necessary to perform semantic tokenization on the semantic information of the retrieval requirement, and then construct a retrieval requirement embedding vector through semantic tokenization; when performing semantic tokenization, it is necessary to perform the tokenization operation in accordance with the semantics, that is, perform semantic tokenization processing on the semantic information of the retrieval requirement according to the semantics, so as to form a tokenization result corresponding to the semantic information of the retrieval requirement, where the tokenization result is to segment the semantic information of the retrieval requirement into a series of tokenized words, and the tokenized words are the basic tokenization units of the semantic information of the retrieval requirement, which are single characters or phrases.
[0057] After obtaining the tokenization result, it is necessary to convert each tokenized word in the tokenization result into a vector in an embedding matrix through a large language model (LLM), and then perform an accumulation process on the vectors to form a retrieval requirement embedding vector, where the embedding matrix is a vector library, and each vector in the vector library corresponds to a specific tokenized word.
[0058] There is a semantic embedding module in the large language model, that is, each tokenized word in the input tokenization result is converted into a vector in an embedding matrix through the semantic embedding module in the large language model; and the accumulation process is also performed in the large language model when performing the accumulation process on the vectors, mainly integrating them into a single vector, which comprehensively incorporates the information input from the entire semantic text information of the retrieval requirement; this comprehensive vector is called a semantic embedding vector, that is, the retrieval requirement embedding vector in the present application; this retrieval requirement embedding vector can capture the overall meaning and context information of the semantic text information of the retrieval requirement.
[0059] Similarity calculation module 303: used to perform vector similarity calculation processing on the retrieval requirement embedding vector and the embedding vectors stored in the index database, and obtain the vector similarity calculation result of the retrieval requirement embedding vector, where the embedding vectors stored in the index database are constructed using the text descriptions stored in the teaching domain data warehouse in the data center; In the specific implementation process of the present invention, calculating the vector similarity between the retrieved demand embedding vector and the embedding vectors stored in the index database to obtain the vector similarity calculation result of the retrieved demand embedding vector includes: using the cosine similarity algorithm to calculate the vector similarity between the retrieved demand embedding vector and the embedding vectors stored in the index database to obtain the first similarity calculation result of the retrieved demand embedding vector; using the Euclidean distance similarity algorithm to calculate the vector similarity between the retrieved demand embedding vector and the embedding vectors stored in the index database to obtain the second similarity calculation result of the retrieved demand embedding vector; performing linear weighting on the first similarity calculation result and the second similarity calculation result to obtain the vector similarity calculation result of the retrieved demand embedding vector.
[0060] Specifically, the data middle platform needs to compare the retrieved demand embedding vectors formed by the input with the embedding vectors stored in the index database, which can be achieved by using similarity. In this application, different metrics can be used to compare the similarity of vectors, and then comprehensive weighting is performed, so that the corresponding similarity can be obtained more accurately; in this application, the cosine similarity algorithm and the Euclidean distance similarity algorithm can be used. In this embodiment, the higher the similarity or the smaller the distance, the stronger the semantic relationship between the vectors.
[0061] That is, using the cosine similarity algorithm to calculate the vector similarity between the retrieved demand embedding vector and the embedding vectors stored in the index database to obtain the first similarity calculation result of the retrieved demand embedding vector; then using the Euclidean distance similarity algorithm to calculate the vector similarity between the retrieved demand embedding vector and the embedding vectors stored in the index database to obtain the second similarity calculation result of the retrieved demand embedding vector; finally, performing linear weighting on the first similarity calculation result and the second similarity calculation result to obtain the vector similarity calculation result of the retrieved demand embedding vector.
[0062] In this embodiment, regarding the index database, embedding vectors are stored in the index database, where the embedding vectors are constructed using the text descriptions stored in the teaching domain data warehouse in the data middle platform; that is, the embedding vectors are generated using the text descriptions corresponding to the data such as application tables, summary tables, and detail tables in the teaching domain data warehouse.
[0063] Selection module 304: used to select the corresponding retrieved embedding vector based on the vector similarity calculation result; In the specific implementation process of the present invention, the selecting of the corresponding retrieval embedded word vector based on the vector similarity calculation result includes: sorting the vector similarity calculation result, and selecting the corresponding retrieval embedded word vector based on the vector similarity sorting result.
[0064] Specifically, the vector similarity calculation results are first sorted, and then the embedded word vectors corresponding to the top few similarities in the vector similarity sorting results are selected as the corresponding retrieval embedded word vectors.
[0065] Data collection module 305: used to construct prompt words based on the retrieval embedded word vector, and call the large language model based on the prompt words to return the teaching field data information that needs to be collected in the data center.
[0066] In the specific implementation process of the present invention, the prompt word is constructed based on the retrieval embedded word vector, and the large language model is called based on the prompt word to return the teaching field data information that needs to be collected in the data center, including: inputting the retrieval embedded word vector into the prompt word construction model, and generating the prompt word based on the prompt word construction model, the prompt word construction model is a pre-trained language model, and during the training process, the input training data is designed, experimented and optimized to guide the model to generate targeted prompt words; calling the large language model based on the prompt word, and inputting the prompt word and the retrieval embedded word vector into the large language model to retrieve and process the teaching field data in the data center to obtain the teaching field data information that needs to be collected.
[0067] Specifically, it is first necessary to construct prompt words by retrieving embedded word vectors, and finally to retrieve and process the teaching field data in the data center through the prompt words and the retrieval embedded word vectors, so as to obtain the teaching field data information that needs to be collected.
[0068] Therefore, it is necessary to input the retrieval embedded word vector into the prompt word construction model, and generate prompt words through the prompt word construction model, wherein the prompt word construction model is a pre-trained language model, and during the training process, the model is guided to generate targeted prompt words by designing, experimenting and optimizing the input training data; after obtaining the prompt word, the corresponding large language model is called through the prompt word, and the prompt word and the retrieval embedded word vector are input into the large language model, and the teaching field data in the data center is retrieved and processed in the large language model, so as to obtain the teaching field data information that needs to be collected.
[0069] That is, the prompt word construction model is a pre-trained speech model, which is a technology that guides the model to generate high-quality, accurate and targeted outputs through design, experimentation and optimization of input prompt words. Prompt engineering is essentially a form of human-computer interaction. The prompt words are the input (instructions) provided to the LLM (large language model). The large model outputs content related to the instructions based on the instructions and its own pre-trained "knowledge". The quality of the output results of the large model is related to the input instructions.
[0070] Based on the reference example of prompt words in Text2Sql, prompt words generally consist of the following elements: Role: Define a role for the big model that matches the target task; its role can be made clear in one sentence (such as "you are a big data engineer"), thereby effectively narrowing the problem domain, reducing ambiguity, and guiding the big model to deepen from "general" to "professional field".
[0071] Constraints: Describe the specific task in detail.
[0072] Context: Provides other background information related to the task (such as historical dialogue, situation, etc.).
[0073] Examples: Examples are very important and are an important reference when generating outputs for large models, and are very helpful for the output results.
[0074] Input: The input information of the task, with a clear "input" mark in the prompt.
[0075] Output: Description of the output format, return results in JSON format, etc.
[0076] According to the Text2Sql solution based on the big model, it is necessary to design prompt words related to the undergraduate teaching domain.
[0077] In an embodiment of the present invention, a data retrieval requirement text is received, and the corresponding retrieval requirement semantic information in the data retrieval requirement text is extracted, and then a retrieval requirement embedded word vector is constructed through the retrieval requirement semantic information; and the corresponding retrieval embedded sub-vector is selected through the similarity with the embedded word vector stored in the index database, and finally, a prompt word is constructed, and the corresponding teaching field data information is collected in the data center through a large language model; by utilizing the semantic embedding and search technology in the large model, the required teaching field data can be accurately retrieved in the data center, thereby improving the efficiency of data collection.
[0078] A computer-readable storage medium provided by an embodiment of the present invention has a computer program stored thereon. When the program is executed by a processor, it implements the teaching field data acquisition method of any one of the above embodiments. Among them, the computer-readable storage medium includes, but is not limited to, any type of disk (including floppy disks, hard disks, optical disks, CD-ROMs, and magneto-optical disks), ROM (Read-Only Memory), RAM (Random Access Memory), EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), flash memory, magnetic cards, or optical cards. That is, the storage device includes any medium that can store or transmit information in a readable form by a device (such as a computer, mobile phone), and can be a read-only memory, a magnetic disk, or an optical disk, etc.
[0079] An embodiment of the present invention also provides a computer application program that runs on a computer and is used to execute the teaching field data acquisition method of any one of the above embodiments.
[0080] In addition, Figure 4 is a schematic structural diagram of an electronic device in an embodiment of the present invention.
[0081] An embodiment of the present invention also provides an electronic device, as Figure 4 shown. The electronic device includes devices such as a processor 402, a memory 403, an input unit 404, and a display unit 405. Those skilled in the art can understand that Figure 4 the structural devices of the electronic device shown do not constitute a limitation on all devices, and may include more or fewer components than shown, or combine certain components. The memory 403 can be used to store the application program 401 and each functional module. The processor 402 runs the application program 401 stored in the memory 403, thereby executing various functional applications and data processing of the device. The memory can be an internal memory or an external memory, or include both an internal memory and an external memory. The internal memory can include a read-only memory (ROM), a programmable ROM (PROM), an electrically programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), flash memory, or a random access memory. The external memory can include a hard disk, a floppy disk, a ZIP disk, a USB flash drive, a magnetic tape, etc. The memory disclosed in the present invention includes, but is not limited to, these types of memories. The memory disclosed in the present invention is only an example and not a limitation.
[0082] The input unit 404 is used to receive the input of signals and the keywords input by the user. The input unit 404 may include a touch panel and other input devices. The touch panel can collect the touch operations of the user on or near it (such as the operations of the user using any suitable object or accessory such as a finger, a stylus, etc. on or near the touch panel), and drive the corresponding connection device according to a pre-set program; the other input devices may include, but are not limited to, a physical keyboard, function keys (such as play control keys, switch keys, etc.), a trackball, a mouse, a joystick, etc. The display unit 405 can be used to display the information input by the user or the information provided to the user and various menus of the terminal device. The display unit 405 can be in the form of a liquid crystal display, an organic light-emitting diode, etc. The processor 402 is the control center of the terminal device, connecting various parts of the entire device through various interfaces and lines, and performing various functions and processing data by running or executing the software programs and / or modules stored in the memory 403, and calling the data stored in the memory.
[0083] As an embodiment, the electronic device includes: one or more processors 402, a memory 403, and one or more application programs 401, wherein the one or more application programs 401 are stored in the memory 403 and are configured to be executed by the one or more processors 402, and the one or more application programs 401 are configured to execute the corresponding teaching field data collection method in any one of the above embodiments.
[0084] In the embodiment of the present invention, by receiving the data retrieval requirement text, extracting the corresponding retrieval requirement semantic information in the data retrieval requirement text, constructing a retrieval requirement embedding word vector through the retrieval requirement semantic information; selecting the corresponding retrieval embedding sub-vector through the similarity of the embedding word vector stored in the index database, and finally constructing a prompt word, and collecting the corresponding teaching field data information in the data center through a large language model; by using the semantic embedding and search technology in the large model, the required teaching field data can be accurately retrieved in the data center, thereby improving the efficiency of data collection.
[0085] In addition, the above has introduced in detail a teaching field data collection method and related devices based on a data center provided by the embodiments of the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A teaching field data collection method based on a data center, characterized in that: The method for data collection applied to the teaching field includes: The data center receives the data retrieval requirement text uploaded by the user based on the retrieval terminal, and extracts the retrieval requirement semantic information in the data retrieval requirement text; Performing semantic word segmentation processing on the retrieval requirement semantic information, and generating a retrieval requirement embedding word vector based on the word segmentation processing result of the retrieval requirement semantic information; Performing vector similarity calculation processing on the embedded word vector of the retrieval requirement and the embedded word vector stored in the index database to obtain the vector similarity calculation result of the embedded word vector of the retrieval requirement, wherein the embedded word vector stored in the index database is constructed by using the text description stored in the teaching domain data warehouse in the data center; Selecting a corresponding retrieval embedding word vector based on the vector similarity calculation result; A prompt word is constructed based on the retrieval embedded word vector, and a large language model is called based on the prompt word to return the teaching field data information that needs to be collected in the data center.
2. The teaching field data collection method according to claim 1 is characterized in that: The data middle station receives the data retrieval requirement text uploaded by the user based on the retrieval terminal, including: After the data center establishes a communication connection with the search terminal, the data center receives the search authority authentication information uploaded by the user based on the search terminal; The data center performs authentication processing on the search authority authentication information, and allocates corresponding search authority to the user on the search terminal based on the authentication result; The data middle station receives a data retrieval requirement text uploaded by the user on the retrieval terminal based on the corresponding retrieval authority allocated.
3. The teaching field data collection method according to claim 1, characterized in that: The retrieval requirement semantic information extracted from the data retrieval requirement text includes: Invoking a corresponding semantic extraction model in the data center, inputting the data retrieval requirement text into the semantic extraction model, and extracting retrieval requirement semantic information from the data retrieval requirement text based on the semantic extraction model; The semantic extraction model is a model obtained by inputting historical data retrieval requirement text and expert-marked semantic content corresponding to the historical data retrieval requirement text into a semantic expression extraction network for training and convergence.
4. The teaching field data collection method according to claim 1, characterized in that: The performing semantic word segmentation processing on the semantic information of the retrieval requirement, and generating a retrieval requirement embedding word vector based on the word segmentation processing result of the semantic information of the retrieval requirement, includes: Performing semantic segmentation processing on the retrieval requirement semantic information in a semantic segmentation manner to form a segmentation processing result corresponding to the retrieval requirement semantic information, wherein the segmentation processing result is to divide the retrieval requirement semantic information into a series of marked segmentations, wherein the marked segmentations are basic segmentation units of the retrieval requirement semantic information, and are single characters or phrases; Based on the large language model, each word segmentation in the word segmentation processing result is converted into a vector in an embedding matrix, and the vectors are accumulated to form a retrieval requirement embedded word vector. The embedding matrix is a vector library, and each vector in the vector library corresponds to a specific marked word segmentation.
5. The teaching field data collection method according to claim 1 is characterized in that: The step of performing vector similarity calculation processing on the embedded word vector of the retrieval requirement and the embedded word vector stored in the index database to obtain the vector similarity calculation result of the embedded word vector of the retrieval requirement includes: Performing vector similarity calculation processing on the retrieval requirement embedded word vector and the embedded word vector stored in the index database using a cosine similarity algorithm to obtain a first similarity calculation result of the retrieval requirement embedded word vector; Using the Euclidean distance similarity algorithm, a vector similarity calculation process is performed on the embedded word vector of the retrieval requirement and the embedded word vector stored in the index database to obtain a second similarity calculation result of the embedded word vector of the retrieval requirement; The first similarity calculation result and the second similarity calculation result are linearly weighted to obtain a vector similarity calculation result of the retrieval requirement embedded word vector.
6. The teaching field data collection method according to claim 1, characterized in that: The selecting a corresponding retrieval embedding word vector based on the vector similarity calculation result includes: The vector similarity calculation results are sorted, and the corresponding retrieval embedding word vector is selected based on the vector similarity sorting results.
7. The teaching field data collection method according to claim 1 is characterized in that: The step of constructing a prompt word based on the search embedding word vector, and calling a large language model based on the prompt word to return the teaching field data information that needs to be collected in the data center, includes: Inputting the search embedding word vector into a prompt word construction model, and generating the prompt word based on the prompt word construction model, wherein the prompt word construction model is a pre-trained language model, and during the training process, the model is guided to generate targeted prompt words by designing, experimenting and optimizing input training data; The large language model is called based on the prompt word, and the prompt word and the retrieval embedded word vector are input into the large language model to perform retrieval processing on the teaching field data in the data center to obtain the teaching field data information that needs to be collected.
8. A teaching field data acquisition device based on a data center, characterized in that: Applied to data collection in the field of teaching, the device comprises: Receiving module: used for the data center to receive the data retrieval requirement text uploaded by the user based on the retrieval terminal, and extract the retrieval requirement semantic information in the data retrieval requirement text; A generation module is used to perform semantic word segmentation processing on the semantic information of the retrieval requirement, and generate a retrieval requirement embedding word vector based on the word segmentation processing result of the semantic information of the retrieval requirement; Similarity calculation module: used for performing vector similarity calculation processing on the embedded word vector of the retrieval requirement and the embedded word vector stored in the index database to obtain the vector similarity calculation result of the embedded word vector of the retrieval requirement, wherein the embedded word vector stored in the index database is constructed by using the text description stored in the teaching domain data warehouse in the data center; A selection module: used for selecting a corresponding retrieval embedding word vector based on the vector similarity calculation result; Data collection module: used to construct prompt words based on the retrieval embedded word vector, and call the large language model based on the prompt words to return the teaching field data information that needs to be collected in the data center.
9. An electronic device, comprising a processor and a memory, characterized in that: The processor runs the computer program or code stored in the memory to implement the teaching field data collection method according to any one of claims 1 to 7.
10. A computer-readable storage medium for storing a computer program or code, characterized in that: When the computer program or code is executed by a processor, the teaching field data collection method as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Knowledge graph-based cue word generation method and device for vector similarity search
CN118210889A
Text search matching method and system based on large language model
CN118260382A
Automatic data entry method and device, electronic equipment and storage medium
CN118395951A
Data retrieval output method based on large language model
CN119621749A
Vector-based document retrieval method and apparatus, computer device, and storage medium
WO2021175005A1