Business English learning method and device based on AI, medium and electronic equipment
Through deep learning model and BERT model, semantic analysis and contextual understanding of business English texts is solved, and the problem of lack of systematic and personalized learning suggestions and insufficient contextual understanding ability in the existing technology is solved, and more efficient and accurate business English learning effect is achieved.
Patent Information
- Application Number
- CN202510007735.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-06
Smart Images

Figure CN119940359A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing and artificial intelligence technology, and specifically to an AI-based business English learning method, device, medium and electronic equipment. Background Art
[0002] With the acceleration of globalization and the continuous development of international trade, business English, as an important tool for international communication, has a growing demand for learning. Traditional business English learning methods often rely on paper textbooks, classroom lectures and manual tutoring, which have problems such as low learning efficiency and lack of personalization. In recent years, with the rapid development of artificial intelligence technology, especially the continuous advancement of natural language processing technology, new solutions have been provided for business English learning.
[0003] At present, there are already some AI-based business English learning products on the market, such as smart dictionaries, online translation tools, and voice recognition systems. However, most of these products can only provide a single vocabulary query, translation, or voice conversion function, and lack systematic learning suggestions and contextual understanding. In addition, some patents also attempt to apply AI technology to business English learning, such as patent number CN109448468A, which is an intelligent teaching system for business English teaching. Although this method can analyze students' learning data through machine learning algorithms, it fails to fully utilize the deep learning model's ability to understand text semantics and contextual relationships, resulting in limited accuracy and personalization of learning suggestions.
[0004] Specifically, the business English learning methods in the existing technology have the following deficiencies: (1) Lack of systematic learning suggestions: Most existing methods can only provide a single vocabulary or grammar explanation, and lack systematic learning suggestions based on students' learning progress and difficulties; (2) Insufficient contextual understanding ability: When processing business English texts, existing methods can often only perform simple vocabulary matching and translation, and lack an in-depth understanding of the overall meaning and context of the text; (3) Low degree of personalization: Existing methods fail to make full use of students' learning data and preferences to provide personalized learning suggestions and exercises. Summary of the invention
[0005] In response to the above problems, the present invention proposes an AI-based business English learning method, device, medium and electronic device, which aims to perform semantic analysis and contextual understanding of business English texts through a deep learning model, generate systematic learning suggestions, vocabulary interpretation, syntactic parsing and contextual understanding results, thereby improving the efficiency and effectiveness of business English learning.
[0006] In order to solve the above technical problems, the technical solution provided by the present invention is: a business English learning method based on AI, comprising the following steps:
[0007] S1. Obtain business English text data and perform cleaning, word segmentation and part-of-speech tagging to generate a text data set;
[0008] S2. Extract features from the text data set to generate a feature vector set, and analyze the feature vector set by the AI model;
[0009] S3. Based on the analysis results, provide users with learning suggestions, vocabulary explanations, syntactic analysis, and contextual understanding feedback.
[0010] Furthermore, S1 further includes the following steps:
[0011] S1.1. Use the database interface to obtain original business English text data from business English learning websites, business emails, business negotiation records, and business English textbooks, expressed as a set D, D = {d 1 ,d 2 ,...,di}, where d i represents the i-th business English text data;
[0012] S1.2. Clean each text data in set D, remove invalid characters, HTML tags, and stop words, and obtain the cleaned text data set D', D' = {d 1′ ,d 2′ ,...,d i′}, where d i ′ represents the i-th text data after cleaning;
[0013] S1.3. Perform word segmentation on each text data in the set D' to obtain a text data set W after word segmentation, W = {w 1 ,w 2 ,...,w m}, where w m Represents the mth word after word segmentation;
[0014] S1.4. Perform part-of-speech tagging on each word in the set W to obtain a text data set P after part-of-speech tagging, P = {(w 1 ,t 1 ),(w 2 ,t 2 ),...,(w m ,t m )}, where (w m ,t m ) represents the mth word and its corresponding part-of-speech tag.
[0015] Furthermore, S1.2 includes the following steps:
[0016] S1.2.1. Define a set of invalid characters C invalid, which contains line breaks, tabs, and extra spaces; traverse d i For each character in C invalid If it is in the text file, replace it with a space or delete it directly to get the business English text data d i,step1 ;
[0017] S1.2.2. Use regular expressions to match and remove business English text data i,step1 HTML tags in the text data to get business English text data i,step2 ;
[0018] S1.2.3. Define the stop word set C stop , traverse d i,step2 For each word in C stop , then delete it, and finally get the cleaned text data set D', D'={d 1′ ,d 2′ ,...,d i′}, where d i′ Represents the i-th text data after cleaning.
[0019] Further, S1.3 includes the following steps:
[0020] S1.3.1. Constructing word segmentation dictionary D diet , the dictionary contains all possible words and their corresponding segmentation results;
[0021] S1.3.2. Initialize an empty list w i , used to store the words after word segmentation; scan d from left to right;
[0022] S1.3.3 Scan from left to right i ′, take out a substring of length l each time;
[0023] S1.3.4. Check whether the substring is in the word segmentation dictionary D die If so, add the substring as a word to w i In, and from d i ', delete the substring; if not, reduce the value of l and repeat step S1.3.3 until a matching word is found or l is reduced to 1; if still not matched, treat the single character as a word,
[0024] S1.3.5. Repeat steps S1.3.3 and S1.3.4 until d i ′ is completely segmented, and the segmented text data set W is obtained, W = {w 1 ,w 2 ,...,w m}, where wm Represents the mth word after word segmentation.
[0025] Further, S1.4 includes the following steps:
[0026] S1.4.1. Use a well-annotated corpus to train a part-of-speech tagging model. The well-annotated corpus contains words and their corresponding part-of-speech tags.
[0027] S1.4.2. Initialize an empty list p i , used to store the results after part-of-speech tagging;
[0028] S1.4.3. Traversing w i For each word in, use the part-of-speech tagging model M pos Tag the part of speech.
[0029] S1.4.4. Add the annotated vocabulary and part-of-speech tags to p i In the text dataset P after part-of-speech tagging, P = {(w 1 ,t 1 ),(w 2 ,t 2 ),...,(w m ,t m )}, where (w m ,t m ) represents the mth word and its corresponding part-of-speech tag.
[0030] Further, S2 includes the following steps:
[0031] S2.1. Input a text data set P, traverse the entire data set P, count the frequency of occurrence of each unique word and part-of-speech pair, and generate the corresponding feature vector value;
[0032] S2.2. For each piece of text data P i The feature vector value of each dimension corresponds to a unique word and part-of-speech pair, and the value of the dimension is the value of the word-part-of-speech pair in the text P i The number of times it appears in the , sum up the eigenvector values to get the eigenvector set;
[0033] S2.3, output a feature vector set, where each element is a feature vector corresponding to a text data in the original data set; divide the data set into 70% training set, 15% validation set and 15% test set according to the proportion;
[0034] S2.4. Load the pre-trained BERT model, map the feature vector to the vocabulary index of the BERT model, and obtain the embedding representation of each word;
[0035] S2.5, the embedded representation is sent to the Transformer encoder layer of the BERT model for processing. The Transformer encoder layer encodes the context of the input sequence through components such as the multi-head self-attention mechanism and the feedforward neural network to obtain the context representation of each position;
[0036] S2.6. Extract the feature vector from the output of the Transformer layer, which contains the semantic information and contextual relationship of the input text. Use the output of the last layer of the BERT model as the feature vector. You can also select the output of its user layer according to specific task requirements.
[0037] Furthermore, S3 includes the following steps:
[0038] S3.1, Generate targeted learning suggestions based on vocabulary and part-of-speech information in the feature vector and their contextual relationship in the text;
[0039] S3.2. Use the semantic information in the feature vector to explain and illustrate the usage of specific words. By calculating the similarity between feature vectors, find the words and examples most relevant to the target words and provide accurate word explanations.
[0040] S3.3. Analyze the dependencies and syntactic structures between words in the feature vector and generate the syntactic parsing results of the text, which can be achieved by adding additional syntactic parsing heads based on the BERT model or using the output of the BERT model as the input of the syntactic parser;
[0041] S3.4. Use the contextual information in the feature vector to understand the overall meaning and context of the text, and perform contextual analysis and understanding of the text by calculating the similarity between feature vectors or using the context encoding ability of the BERT model;
[0042] S3.5. Output the generated learning suggestions, vocabulary explanations, syntactic parsing and context understanding results to the user.
[0043] Furthermore, an AI-based business English learning device is provided, which is characterized by comprising:
[0044] The text data preprocessing module is connected to the business English data source through the database interface to realize automatic data acquisition and preprocessing; it obtains business English text data from multiple channels, and performs cleaning, word segmentation and part-of-speech tagging to generate a text data set; it receives the original text data sent by the data source, and after processing, outputs the cleaned text data set and word segmentation and part-of-speech tagging results.
[0045] The feature extraction and analysis module extracts features from the preprocessed text dataset based on the BERT model, generates a feature vector set, and uses the BERT model to analyze the feature vector to extract semantic information and contextual relationships; receives the text dataset output by the preprocessing module, and outputs a feature vector set and the context representation of the BERT model;
[0046] The learning suggestion generation module generates targeted learning suggestions based on the vocabulary and part-of-speech information in the feature vector and the contextual relationship; analyzes the difficulties and error points of the user in the learning process and provides personalized learning suggestions and exercises; receives the feature vector and context representation output by the feature analysis module and outputs learning suggestions;
[0047] The vocabulary explanation module uses the semantic information in the feature vector to explain and explain the usage of specific vocabulary; by calculating the similarity between feature vectors, it finds the vocabulary and examples most relevant to the target vocabulary and provides accurate vocabulary explanations and usage instructions; it receives the feature vectors output by the feature analysis module and outputs vocabulary explanations and examples.
[0048] The syntactic parsing module, based on the BERT model, adds an additional syntactic parsing head or uses the output of the BERT model as the input of the syntactic parser; analyzes the dependency relationship and syntactic structure between words in the feature vector to generate the syntactic parsing result of the text; receives the feature vector output by the feature analysis module and outputs the syntactic parsing result.
[0049] The context understanding module uses the context information in the feature vector to understand the overall meaning and context of the text; performs context analysis and understanding of the text by calculating the similarity between feature vectors or using the context encoding capability of the BERT model; receives the feature vector and context representation output by the feature analysis module, and outputs the context understanding result.
[0050] The user interaction module provides users with an intuitive and convenient learning interface, enabling users to easily obtain learning suggestions, query vocabulary explanations, view syntactic parsing results and provide feedback on contextual understanding results; receives learning suggestions, vocabulary explanations, syntactic parsing and contextual understanding results output by each functional module, displays them to users, and receives user feedback and interaction.
[0051] Furthermore, a computer-readable storage medium is provided, characterized in that: the storage medium stores a computer program, and when the computer program is executed by a processor, the method of any one of claims 1 to 7 is implemented.
[0052] Furthermore, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the method of any one of claims 1 to 7 is implemented when the processor executes the program.
[0053] Compared with the prior art, the advantages of the present invention are: (1) Systematic learning suggestions: The present invention uses a deep learning model to perform semantic analysis and contextual understanding of business English texts, and can generate systematic learning suggestions for students' learning progress and difficulties, thereby improving learning efficiency; (2) Powerful contextual understanding ability: The present invention uses the BERT model to contextually encode the text, which can deeply understand the overall meaning and context of the text and provide more accurate vocabulary interpretation and syntactic analysis results; (3) High degree of personalization: The present invention provides personalized learning suggestions and exercises based on students' learning data and preferences to meet the learning needs of different students. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 This is a schematic diagram of the steps of the AI-based business English learning method.
[0055] Figure 2 This is a schematic diagram of the composition of an AI-based business English learning device. DETAILED DESCRIPTION
[0056] The present invention is further described in detail below in conjunction with the accompanying drawings.
[0057] Embodiment 1
[0058] like Figure 1 and Figure 2 As shown, this embodiment provides an AI-based business English learning method, comprising the following steps:
[0059] S1. Obtain business English text data and perform cleaning, word segmentation and part-of-speech tagging to generate a text data set;
[0060] S2. Extract features from the text data set to generate a feature vector set, and analyze the feature vector set by the AI model;
[0061] S3. Based on the analysis results, provide users with learning suggestions, vocabulary explanations, syntactic analysis, and contextual understanding feedback.
[0062] Among them, step S1 also includes the specific steps of using the database interface to obtain business English text data from multiple channels, and performing cleaning, word segmentation and part-of-speech tagging; step S2 also includes the specific steps of feature extraction and analysis based on the BERT model; step S3 also includes the specific steps of generating learning suggestions, vocabulary interpretation, syntactic parsing and context understanding results based on feature vectors.
[0063] In addition, the present invention also provides an AI-based business English learning device, including a text data preprocessing module, a feature extraction and analysis module, a learning suggestion generation module, a vocabulary interpretation module, a syntax parsing module, a context understanding module and a user interaction module. The modules work together to achieve intelligent processing of business English texts and generation of learning suggestions.
[0064] Specifically, in S1.1 of this embodiment, the database interface is used to obtain original business English text data from business English learning websites, business emails, business negotiation records, and business English textbooks, which are represented as a set D, where D = {d 1 , d 2 , .., d i}, where d i represents the i-th business English text data;
[0065] In S1.2 of this embodiment, each text data in the set D is cleaned to remove invalid characters, HTML tags, and stop words to obtain a cleaned text data set D′, D={d 1′ , d 2′ , .., d i′}, where d i′ Represents the i-th text data after cleaning, specifically:
[0066] S1.2.1. Define a set of invalid characters C invalid , which contains line breaks, tabs, and extra spaces; traverse d i For each character in C invalid If it is in the text file, replace it with a space or delete it directly to get the business English text data d i,step1 ,formula:
[0067] d i,step1 =replace_invalid_chars(d i , C invlid )
[0068] Among them, replace_invalid_chars is a function used to replace or delete invalid characters.
[0069] S1.2.2. Use regular expressions to match and remove business English text data i,step1 HTML tags in the text data to get business English text data i,step2 ,formula:
[0070] d i,step2 =remove_html_tags(d i,step1 )
[0071] Among them, remove_html_tags is a function that uses regular expressions to match and remove HTML tags.
[0072] S1.2.3. Define the stop word set C stop , traverse d i,step2 For each word in C stop , then delete it, and finally get the cleaned text data set D′, D′={d 1′ , d 2′ , ..., d i′}, where d i′ Represents the i-th text data after cleaning, formula:
[0073] d i =remove_stop_words(d i,step2 , C stop )
[0074] Among them, remove_stop_words is a function used to remove stop words. Before deleting stop words, the text is first segmented into vocabulary sequences.
[0075] In S1.3 of this embodiment, each text data in the set D′ is segmented to obtain a text data set W after segmentation, where W={w 1 , w 2 , .., w m}, where w m It means the mth word after word segmentation. Specifically:
[0076] S1.3.1. Constructing word segmentation dictionary D diet , the dictionary contains all possible words and their corresponding segmentation results, formula:
[0077] D diet =build_dictionary(corpus)
[0078] Among them, corpus represents the corpus used to train the word segmentation dictionary, and build_dictionary is a function used to build the word segmentation dictionary D from the corpus. die .
[0079] S1.3.2. Initialize an empty list w i , used to store the words after word segmentation; scan d from left to right;
[0080] S1.3.3 Scan from left to right i ′, take out a substring of length 1 each time;
[0081] S1.3.4. Check whether the substring is in the word segmentation dictionary D die If so, add the substring as a word to w i In, and from d i ', delete the substring; if not, reduce the value of 1 and repeat step S1.3.3 until a matching word is found or 1 is reduced to 1; if still not matched, treat the single character as a word,
[0082] S1.3.5. Repeat steps S1.3.3 and S1.3.4 until d i ′ is completely segmented, and the segmented text data set W is obtained, W = {w 1 , w 2 , ..., w m}, where w m Represents the mth word after word segmentation.
[0083] formula:
[0084] w i =segment_text(d′ i , D dict )
[0085] Among them, segment_text is a function used to segment text i ' Perform word segmentation and return the word list w after word segmentation i .
[0086] In S1.4 of this embodiment, each word in the set W is tagged with a part of speech to obtain a text data set P after the part of speech tagging, where P = {(w 1 , t 1 ), (w 2 ,t 2 ), .., (w m , t m )}, where (w m , t m ) represents the mth word and its corresponding part-of-speech tag, specifically:
[0087] S1.4.1. Use the annotated corpus to train the part-of-speech tagging model. The annotated corpus contains words and their corresponding part-of-speech tags. The formula is:
[0088] M POS =train_POS_model(anmotated_corpus)
[0089] Where annotated_corpus represents the annotated corpus, and train_POS_model is a function used to train the part-of-speech tagging model M from the annotated corpus. pos .
[0090] S1.4.2. Initialize an empty list p i , used to store the results after part-of-speech tagging;
[0091] S1.4.3. Traversing w i For each word in, use the part-of-speech tagging model M pos Tag the part of speech.
[0092] s1.4.4. Add the annotated vocabulary and part-of-speech tags to p i In the text dataset P after part-of-speech tagging, P = {(w 1 , t 1 ), (w 2 ,t 2 ),...,(w m , t m )}, where (w m , t m ) represents the mth word and its corresponding part-of-speech tag.
[0093] formula:
[0094] p i =[annotate_word(w,M POS )for w∈w i ]
[0095] Among them, annotate_word is a function used to perform part-of-speech tagging on the word w and return the annotated word (word + part-of-speech tag).
[0096] Embodiment 2
[0097] like Figure 1 and Figure 2 As shown, this embodiment provides a specific implementation of S2:
[0098] S2.1. Input the text data set P, traverse the entire data set P, count the frequency of occurrence of each unique word and part-of-speech pair, and generate the corresponding feature vector value. The following is part of the execution code:
[0099] word_pos_freq={}forp in P:
[0100] for(word,pos)in p:
[0101] if(word,pos)in word_pos_freq:
[0102] word_pos_freq[(word,pos)]+=1else:
[0103] word_pos_freq[(word,pos)]=1
[0104] S2.2. For each piece of text data P i The feature vector value of each dimension corresponds to a unique word and part-of-speech pair, and the value of the dimension is the value of the word-part-of-speech pair in the text P i The number of times it appears in the , sum up the eigenvector values to get the eigenvector set;
[0105] S2.3, output a feature vector set, where each element is a feature vector corresponding to a text data in the original data set; divide the data set into 70% training set, 15% validation set and 15% test set according to the proportion;
[0106] S2.4. Load the pre-trained BERT model, map the feature vector to the vocabulary index of the BERT model, and obtain the embedding representation of each word;
[0107] S2.5, the embedded representation is sent to the Transformer encoder layer of the BERT model for processing. The Transformer encoder layer encodes the context of the input sequence through components such as the multi-head self-attention mechanism and the feedforward neural network to obtain the context representation of each position;
[0108] S2.6. Extract the feature vector from the output of the Transformer layer, which contains the semantic information and contextual relationship of the input text. Use the output of the last layer of the BERT model as the feature vector. You can also select the output of its user layer according to specific task requirements.
[0109] Embodiment 3
[0110] like Figure 1 and Figure 2As shown, this embodiment provides a specific implementation of S2: In step S2, the feature extraction and analysis module is used to process the preprocessed text data set. The module first receives the text data set output by the preprocessing module, which has been cleaned, segmented and part-of-speech tagged. Then, the module traverses the entire data set, counts the frequency of occurrence of each unique vocabulary and part-of-speech pair, and generates the corresponding feature vector value. For the feature vector value of each text data, the module summarizes it into a feature vector set, where each dimension corresponds to a unique vocabulary and part-of-speech pair, and the value of the dimension is the number of times the vocabulary-part-of-speech pair appears in the text. The data set is divided into 70% training set, 15% validation set and 15% test set according to the proportion to ensure the accuracy and generalization ability of the model. After loading the pre-trained BERT model, the module maps the feature vector to the vocabulary index of the BERT model to obtain the embedded representation of each vocabulary. These embedded representations are sent to the Transformer encoder layer of the BERT model for processing. Through components such as the multi-head self-attention mechanism and the feedforward neural network, the input sequence is contextually encoded to obtain the context representation of each position. Finally, feature vectors are extracted from the output of the Transformer layer, which contain the semantic information and contextual relationships of the input text, providing strong support for subsequent learning suggestion generation.
[0111] Embodiment 4
[0112] like Figure 1 and Figure 2 As shown, this embodiment provides a specific implementation method of S2: improving business English ability through AI technology. In step S2, a large amount of business English text data is collected and preprocessed, including cleaning, word segmentation and part-of-speech tagging; a feature extraction tool is used to extract features from the preprocessed text data set to generate a feature vector set. This tool is based on the BERT model and can accurately extract semantic information and contextual relationships in the text. In order to verify the effect of feature extraction, the data set is divided into a training set, a validation set and a test set. After obtaining the feature vector set, these feature vectors are used for analysis to find common vocabulary, phrases and syntactic structures in business English. This information provides an important basis for subsequent learning suggestion generation and vocabulary interpretation.
[0113] Embodiment 5
[0114] like Figure 1 and Figure 2 As shown, this embodiment provides a specific implementation of S3: in step S3, the learning suggestion generation module, the vocabulary interpretation module, the syntax parsing module and the context understanding module are used to provide users with personalized learning suggestions, vocabulary interpretation, syntax parsing and context understanding results.
[0115] The learning suggestion generation module generates targeted learning suggestions for users based on the vocabulary and part-of-speech information in the feature vectors, as well as the contextual relationship. These suggestions include recommended learning materials, exercise questions, and learning strategies. The vocabulary interpretation module uses the semantic information in the feature vectors to explain and explain the usage of specific vocabulary. By calculating the similarity between feature vectors, the module finds the vocabulary and example sentences most relevant to the target vocabulary, and provides users with accurate vocabulary interpretation and usage instructions. Based on the BERT model, the syntactic parsing module analyzes the dependencies and syntactic structures between the vocabulary in the feature vectors, and generates syntactic parsing results for the text. These results help users better understand the structure and grammatical rules of sentences. The context understanding module uses the contextual information in the feature vectors to perform contextual analysis and understanding of the text. By calculating the similarity between feature vectors or using the context encoding ability of the BERT model, the module derives the overall meaning and contextual information of the text, and provides users with corresponding interpretations and instructions.
[0116] Embodiment 6
[0117] like Figure 1 and Figure 2 As shown, this embodiment provides a specific implementation method of S3: using an AI-based business English learning method to improve business English ability. In step S3, the user interacts with the system through the user interaction module, and the user enters his or her learning goals and interests through the user interaction module. The system generates a personalized learning plan and learning suggestions for the user based on this information. These suggestions include recommended learning resources, learning paths, and learning progress. During the learning process, the user encounters an unfamiliar business vocabulary. Through the vocabulary interpretation module, the user enters the feature vector of the vocabulary (or directly queries the vocabulary in the system), and the system provides the user with detailed vocabulary explanations and usage instructions, including common collocations, examples, and synonyms of the vocabulary. In order to better understand a complex business English sentence, the user uses the syntactic parsing module to perform syntactic parsing on the sentence. The system generates a syntactic tree and dependency graph of the sentence for the user, and explains the grammatical relationship and semantic connection between the various components in the sentence. When reading a business English article, the user encountered an incomprehensible sentence. Through the context understanding module, the user inputs the feature vector of the sentence (or directly selects the sentence in the system), and the system provides the user with the context analysis and understanding results of the sentence, including the overall meaning of the sentence, context information, and related background knowledge. This information helps the user better understand the content and intention of the article.
[0118] The above shows and describes the basic principles and main features of the present invention and the advantages of the invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention to be protected. The scope of protection of the present invention is defined by the attached claims and their equivalents.
Claims
1. AI-based business English learning method, characterized by The following steps are involved: S1. Obtain business English text data and perform cleaning, word segmentation and part-of-speech tagging to generate a text data set; S2. Extract features from the text data set to generate a feature vector set, and analyze the feature vector set by the AI model; S3. Based on the analysis results, provide users with learning suggestions, vocabulary explanations, syntactic analysis, and contextual understanding feedback.
2. The AI-based business English learning method according to claim 1, characterized in that The S1 further comprises the following steps: S1.
1. Use the database interface to obtain original business English text data from business English learning websites, business emails, business negotiation records, and business English textbooks, expressed as a set D, D = {d1, d2, ..., di}, where d i represents the i-th business English text data; S1.
2. Clean each text data in set D, remove invalid characters, HTML tags, and stop words, and obtain the cleaned text data set D', D' = {d 1′ ,d 2′ ,...,d i′ }, where d i ′ represents the i-th text data after cleaning; S1.
3. Perform word segmentation on each text data in the set D' to obtain a word segmented text data set W, where W = {w1, w2, ..., w m }, where w m Represents the mth word after word segmentation; S1.
4. Perform part-of-speech tagging on each word in the set W to obtain a text data set P after part-of-speech tagging, where P = {(w1, t1), (w2, t2), ..., (w m ,t m )}, where (w m ,t m ) represents the mth word and its corresponding part-of-speech tag.
3. The AI-based business English learning method according to claim 2 is characterized in that The S1.2 comprises the following steps: S1.2.
1. Define a set of invalid characters C invalid , which contains line breaks, tabs, and extra spaces; traverse d i For each character in C invalid If it is in the text file, replace it with a space or delete it directly to get the business English text data d i,step1 ; S1.2.
2. Use regular expressions to match and remove business English text data i,step1 HTML tags in the text data to get business English text data i,step2 ; S1.2.
3. Define the stop word set C stop , traverse d i,step2 For each word in C stop , then delete it, and finally get the cleaned text data set D', D'={d 1′ ,d 2′ ,...,d i′ }, where d i ′ represents the i-th text data after cleaning.
4. The AI-based business English learning method according to claim 2 is characterized in that The S1.3 comprises the following steps: S1.3.
1. Constructing word segmentation dictionary D diet , the dictionary contains all possible words and their corresponding segmentation results; S1.3.
2. Initialize an empty list w i , used to store the words after word segmentation; scan d from left to right; S1.3.3 Scan from left to right i ′, take out a substring of length l each time; S1.3.
4. Check whether the substring is in the word segmentation dictionary D die If so, add the substring as a word to w i In, and from d i ', delete the substring; if not, reduce the value of l and repeat step S1.3.3 until a matching word is found or l is reduced to 1; if still not matched, treat the single character as a word, S1.3.
5. Repeat steps S1.3.3 and S1.3.4 until d i ′ is completely segmented, and the segmented text data set W is obtained, W = {w1,w2,...,w m }, where w m Represents the mth word after word segmentation.
5. The AI-based business English learning method according to claim 2 is characterized in that The S1.4 comprises the following steps: S1.4.
1. Use a well-annotated corpus to train a part-of-speech tagging model. The well-annotated corpus contains words and their corresponding part-of-speech tags. S1.4.
2. Initialize an empty list p i , used to store the results after part-of-speech tagging; S1.4.
3. Traversing w i For each word in, use the part-of-speech tagging model M pos Tag the part of speech. S1.4.
4. Add the annotated vocabulary and part-of-speech tags to p i In the above example, we obtain the text data set P after part-of-speech tagging, P = {(w1, t1), (w2, t2), ..., (w m ,t m )}, where (w m ,t m ) represents the mth word and its corresponding part-of-speech tag.
6. The AI-based business English learning method according to claim 1, characterized in that The S2 comprises the following steps: S2.
1. Input a text data set P, traverse the entire data set P, count the frequency of occurrence of each unique word and part-of-speech pair, and generate the corresponding feature vector value; S2.
2. For each piece of text data P i The feature vector value of each dimension corresponds to a unique word and part-of-speech pair, and the value of the dimension is the value of the word-part-of-speech pair in the text P i The number of times it appears in the , sum up the eigenvector values to get the eigenvector set; S2.3, output a feature vector set, where each element is a feature vector corresponding to a text data in the original data set; divide the data set into 70% training set, 15% validation set and 15% test set according to the proportion; S2.
4. Load the pre-trained BERT model, map the feature vector to the vocabulary index of the BERT model, and obtain the embedding representation of each word; S2.5, the embedded representation is sent to the Transformer encoder layer of the BERT model for processing. The Transformer encoder layer encodes the context of the input sequence through components such as the multi-head self-attention mechanism and the feedforward neural network to obtain the context representation of each position; S2.
6. Extract the feature vector from the output of the Transformer layer, which contains the semantic information and contextual relationship of the input text. Use the output of the last layer of the BERT model as the feature vector. You can also select the output of other layers according to specific task requirements.
7. The AI-based business English learning method according to claim 1, characterized in that The S3 comprises the following steps: S3.1, Generate targeted learning suggestions based on vocabulary and part-of-speech information in the feature vector and their contextual relationship in the text; S3.
2. Use the semantic information in the feature vector to explain and illustrate the usage of specific words. By calculating the similarity between feature vectors, find the words and examples most relevant to the target words and provide accurate word explanations. S3.
3. Analyze the dependencies and syntactic structures between words in the feature vector and generate the syntactic parsing results of the text, which can be achieved by adding additional syntactic parsing heads based on the BERT model or using the output of the BERT model as the input of the syntactic parser; S3.
4. Use the contextual information in the feature vector to understand the overall meaning and context of the text, and perform contextual analysis and understanding of the text by calculating the similarity between feature vectors or using the context encoding ability of the BERT model; S3.
5. Output the generated learning suggestions, vocabulary explanations, syntactic parsing and context understanding results to the user.
8. The AI-based business English learning device is characterized by include: The text data preprocessing module connects to the business English data source through the database interface to achieve automatic data acquisition and preprocessing; Acquire business English text data from multiple channels, and perform cleaning, word segmentation and part-of-speech tagging to generate text datasets; receive raw text data sent by the data source, process and output the cleaned text dataset and word segmentation and part-of-speech tagging results. The feature extraction and analysis module extracts features from the preprocessed text dataset based on the BERT model, generates a feature vector set, and uses the BERT model to analyze the feature vector to extract semantic information and contextual relationships; receives the text dataset output by the preprocessing module, and outputs a feature vector set and the context representation of the BERT model; The learning suggestion generation module generates targeted learning suggestions based on the vocabulary and part-of-speech information in the feature vector and the context relationship; Analyze the difficulties and mistakes that users encounter during the learning process and provide personalized learning suggestions and exercises; Receive the feature vector and context representation output by the feature analysis module and output learning suggestions; The vocabulary explanation module uses the semantic information in the feature vector to explain and explain the usage of specific vocabulary. By calculating the similarity between feature vectors, it finds the most relevant vocabulary and example sentences for the target vocabulary and provides accurate vocabulary explanation and usage instructions. Receive the feature vector output by the feature analysis module and output vocabulary explanations and example sentences. The syntactic parsing module adds an additional syntactic parsing head based on the BERT model or uses the output of the BERT model as the input of the syntactic parser; Analyze the dependency and syntactic structure between words in the feature vector to generate the syntactic parsing results of the text; Receive the feature vector output by the feature analysis module and output the syntactic parsing result. The context understanding module uses the context information in the feature vector to understand the overall meaning and context of the text; it performs context analysis and understanding of the text by calculating the similarity between feature vectors or using the context encoding ability of the BERT model; Receive the feature vector and context representation output by the feature analysis module, and output the context understanding result. The user interaction module provides users with an intuitive and convenient learning interface, enabling users to easily obtain learning suggestions, query vocabulary explanations, view syntactic parsing results and provide feedback on contextual understanding results; receives learning suggestions, vocabulary explanations, syntactic parsing and contextual understanding results output by each functional module, displays them to users, and receives user feedback and interaction.
9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Intelligent teaching system for business English teaching
CN109448468A
Cited By
Intelligent analysis and text optimization method and system for external financial and economic words
CN121543600A
Method and system for intelligent analysis and text optimization of foreign financial discourse
CN121543600B