An AIGC-based foreign language training content generation method

CN118964573BActive Publication Date: 2026-09-08YUANYU DIGITAL MANUFACTURING (XIAMEN) TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411064434.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-05
Publication Date
2026-09-08
Estimated Expiration
2044-08-05

AI Technical Summary

Technical Problem

[0003]语言学习需要在合适场景或情景下进行训练,刚刚学习的单词,更需要通过口语练习加深印象和深入掌握,然而,在学习过程中往往仅仅是单纯的朗读,一方面学习过程枯燥,另一方面难以通过单纯朗读加强记忆,最终导致学习者的接受效果有限

Benefits of technology

本发明能够根据学习者的能力水平选定学习内容,并基于学习内容生成外语场景图像,能够帮助学习者在外语学习过程中进行理解和记忆。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118964573B_ABST
    Figure CN118964573B_ABST
Patent Text Reader

Abstract

The application discloses a foreign language training content generation method based on AIGC, and comprises the following steps: constructing a corpus database and an original material library; performing scene classification on the corpus database; obtaining a material feature library and performing scene classification; obtaining matched corpus sub-databases and material feature sub-databases; obtaining a scene image based on content generation; performing scene recognition and matching degree calculation on the scene image; displaying the scene image and performing oral training; and collecting pronunciation error information for correction training. The application can select learning content according to the ability level of learners, generate a foreign language scene image based on the learning content, and help the learners to understand and remember in the foreign language learning process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method for generating foreign language training content based on AIGC. Background Technology

[0002] AIGC (Artificial Intelligence Generated Content) is an abbreviation for content generated by artificial intelligence. It can be used to create various types of content, including text, images, pictures, and music. This technology can be used in fields such as movies, games, advertising, and animation, and can save production costs and time.

[0003] Language learning requires training in appropriate scenarios or contexts. Newly learned words need to be reinforced and mastered through oral practice. However, the learning process often involves simply reading aloud, which is not only tedious but also makes it difficult to strengthen memory through mere reading, ultimately resulting in limited learning outcomes for the learner. Summary of the Invention

[0004] To address the above problems, this invention provides a method for generating foreign language training content based on AIGC.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: A method for generating foreign language training content based on AIGC includes the following steps: S1. Construct a corpus database and a raw material library, wherein the corpus database contains corpus data for foreign language training, and the raw material library contains raw material data carrying text information and image information; S2. For the corpus database, classify the corpus data according to the pre-defined scene categories, and divide the corpus database into multiple corpus sub-databases based on the scene categories; S3. For the original material library, a material feature database is obtained using a question-answering model. The material feature database includes at least feature IDs, semantic features, and corresponding image regions. Based on the semantic features of the material feature database, scene classification is performed according to a pre-defined scene category, and the material feature database is divided into multiple material feature sub-databases based on the scene category. S4. Obtain the scenario category information for the foreign language training to be conducted, match the corpus sub-database and material feature sub-database based on the scenario category information, and randomly extract a set of target corpus data for the currently matched corpus sub-database. S5. For each target corpus data, use a question-answering model to parse it and obtain a corresponding set of training semantic features P. Where P is the set of training semantic features, p iLet be the i-th training semantic feature, and m be the total number of training semantic features. Based on the training semantic features, search for matching semantic features and corresponding image region material Q in the corresponding material feature sub-database. Where Q is the set of training semantic features, Let j be the image region material, and n be the total number of image region materials. A set of scene images D is generated based on the AIGC content generation method. Where D is the set of scene images, D k Let s be the k-th scene image, and s be the total number of scene images; S6. Perform scene recognition on the scene image generated in step S5 to determine whether it matches the scene category information. If not, discard it; if it does, calculate the matching score between the scene image and the trained semantic features through semantic relevance analysis and sort them. Keep the scene image with the highest matching score as the target scene image. Semantic relevance analysis is implemented through the following methods: The question-answering model is used to parse the scene image to obtain a set of scene semantic features D. kr , , where D kr For the k-th scene image D k A set of parsed scene semantic features. For the Dth kr For the k-th scene image D k The j-th scene semantic feature is parsed, and r is the total number of scene semantic features parsed; The obtained set of scene semantic features D kr Perform feature mapping with a corresponding set of trained semantic features P; Calculate the feature similarity value between the scene semantic features and the training semantic features, and use the feature similarity value as the matching score. The feature similarity value is calculated using the following formula: , where W(D) k P) represents the k-th scene image D. k The feature similarity value between the scene semantic features and the training semantic features P, T ji for With p i The distance weights between the two features, c(j,i) are... With p i The distance between two features; S7. Display the target corpus data and the corresponding target scene image through the display device. The trainee performs oral training. The trainee's voice data is obtained through the sound receiving device. The voice data is analyzed to obtain the oral training score. If the oral training score is greater than the preset score threshold, the next target corpus data and the corresponding target scene image are displayed through the display device. If the oral training score is less than the preset score threshold, the trainee is prompted to retrain.

[0006] S8. Collect pronunciation error information of trainees during oral training, obtain a correction training dataset based on the foreign language words corresponding to the pronunciation error information, extract a set of matching target corpus data from the corpus sub-database based on the correction training dataset, and execute step S5.

[0007] Preferably, the scene recognition in step S6 is achieved through the following method: A scene recognition model is constructed, sample images are obtained, background sample images are obtained by performing background recognition on the sample images, the background sample images are input into the scene recognition model to obtain a first recognition result, the training sample images are input into the scene recognition model to obtain a second recognition result, the target model loss value is obtained based on the difference between the first recognition result and the second recognition result, and the model parameters of the scene recognition model are adjusted based on the target model loss value to obtain a trained scene recognition model. The scene image is input into the trained scene recognition model to obtain the scene recognition result.

[0008] Preferably, the scene recognition in step S6 is achieved through the following method: Semantic information of scene images is obtained using a question-answering model; Based on the obtained semantic information, a text recognition algorithm is used to obtain scene recognition results.

[0009] Preferably, the acquisition of the material feature database using the question-answering model in step S3 is achieved through the following method: Construct a question-answering model, which includes an autocorrelation attention module and a cross-modal interaction attention module; Autocorrelation attention modules are used to extract features from both text and image information. Autocorrelation learning captures features between image regions and between text characters to assess semantic autocorrelation and update features. Given text information, word embedding encoding is used for representation, and word vector features are learned and extracted. For each given text information, the text feature Y is obtained, represented as: , where t j Let be the feature vector of the j-th word, and m be the total number of region features; Given image information T, a feature extraction network is trained to obtain image region features X, which are represented as follows: , where r i Let be the i-th regional feature, and n be the total number of regional features; The semantic associations between images and text are learned and features are updated using a cross-modal interactive attention module; A cascaded approach is used to stack attention layers, and the image region features and text features are updated at multiple levels to obtain the final image region features and text features with clear semantic representation.

[0010] Preferably, the feature extraction of the text information adopts a gated recurrent unit approach, and the feature extraction of the image information utilizes a Faster R-CNN network.

[0011] Preferably, the scenario category is the environment or situation in which the language expression takes place.

[0012] Preferably, in step S2, the corpus data is classified into scenarios according to pre-defined scenario categories, specifically as follows: Build a category rule library based on predefined scene categories; The pre-labeled corpus data samples are input into the scene classification model to obtain the trained scene classification model; The corpus data is input into the trained scene classification model to obtain category labels, and the corpus data is then labeled. Based on the category labels of the corpus data, the corresponding scene category is matched in the category rule base.

[0013] Preferably, the pronunciation error information is a missed reading, misreading, or inaccurate tone.

[0014] By adopting the above technical solution, the present invention has the following advantages compared with the prior art: This invention can select learning content according to the learner's ability level and generate foreign language scene images based on the learning content, which can help learners understand and memorize foreign languages ​​during the learning process. Attached Figure Description

[0015] Figure 1 This is a schematic diagram of the process of the present invention. Detailed Implementation

[0016] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0017] AIGC, or Artificial Intelligence Generated Content, is a technology that utilizes artificial intelligence algorithms to generate creative and high-quality content. Through training models and learning from large amounts of data, AIGC can generate relevant content based on input conditions or guidance. For example, by inputting keywords, descriptions, or samples, AIGC can generate matching articles, images, audio, etc.

[0018] Visual question answering (VQA) aims to automatically answer a natural language question related to the content of a given image. Its task involves research in the emerging interdisciplinary field of computer vision, natural language processing, and artificial intelligence. By simulating real-world scenarios, visual question answering tasks have very wide and meaningful applications in practice.

[0019] Please see Figure 1 This invention discloses a method for generating foreign language training content based on AIGC, comprising the following steps: S1. Construct a corpus database and a raw material library. The corpus database contains corpus data for foreign language training, and the raw material library contains raw material data carrying text and image information. In this embodiment, the corpus data is a foreign language sentence, and the raw material data is an image with foreign language text, or it can be a frame image of a video with foreign language letters.

[0020] S2. For the corpus database, classify the corpus data according to pre-defined scene categories, and divide the corpus database into multiple sub-databases based on scene categories. This step specifically achieves scene classification through the following methods: Build a category rule base based on pre-defined scene categories; input pre-labeled corpus data samples into the scene classification model to obtain a trained scene classification model; input corpus data into the trained scene classification model to obtain category labels, and label the corpus data accordingly; match the corresponding scene category in the category rule base based on the corpus data category labels. In this embodiment, the scene category refers to the environment or situation in which the language expression takes place. Environments include, but are not limited to, cafes, restaurants, stations, riverbanks, etc., and situations include, but are not limited to, work, parties, travel, shopping, etc.

[0021] S3. For the original material library, a material feature database is obtained using a question-answering model. The material feature database includes at least feature IDs, semantic features, and corresponding image regions. Based on the semantic features of the material feature database, scene classification is performed according to pre-defined scene categories, and the material feature database is further divided into multiple material feature sub-databases based on scene categories. The method for obtaining the material feature database using a question-answering model in this step is as follows: A question-answering model is constructed, which includes an autocorrelation attention module and a cross-modal interaction attention module.

[0022] Autocorrelation attention modules are used to extract features from both text and image information. Autocorrelation learning captures features between image regions and between text characters to assess semantic autocorrelation and update features. Given text information, word embedding encoding is used for representation, and word vector features are learned and extracted. For each given text information, the text feature Y is obtained, represented as: , where t j Let be the feature vector of the j-th word, and m be the total number of region features; Given image information T, a feature extraction network is trained to obtain image region features X, which are represented as follows: , where r i Let be the i-th regional feature, and n be the total number of regional features.

[0023] This paper utilizes a cross-modal interactive attention module to learn the semantic relationships between images and text and update features. Specifically, based on image region features X and text features Y, attention weight matrices for images and text are obtained to establish autocorrelation between the two modalities. Image region features X are used as input to an image-guided text self-attention model, and the inner product of image region features and text features is calculated. Similarly, text features Y are used as input to a text-guided image self-attention model, and the inner product of text features and image region features is calculated. The inner product results are normalized to obtain two bidirectional cross-modal weight matrices between images and text. The text features and image region features are then weighted and updated using these cross-modal weight matrices.

[0024] A cascaded approach is used to stack attention layers, and the image region features and text features are updated at multiple levels to obtain the final image region features and text features with clear semantic representation.

[0025] In this embodiment, feature extraction of text information adopts the gated recurrent unit method, and feature extraction of image information utilizes the Faster R-CNN network.

[0026] S4. Obtain the scenario category information for the foreign language training to be conducted, match the corpus sub-database and material feature sub-database based on the scenario category information, and randomly extract a set of target corpus data for the currently matched corpus sub-database.

[0027] S5. For each target corpus data, use a question-answering model to parse it and obtain a corresponding set of training semantic features P. Where P is the set of training semantic features, p iLet be the i-th training semantic feature, and m be the total number of training semantic features; Based on the trained semantic features, the system searches for matching semantic features and corresponding image region material Q in the corresponding material feature sub-database. Where Q is the set of training semantic features, Let j be the image region material, and n be the total number of image region materials; A set of scene images D is generated based on the AIGC content generation method. Where D is the set of scene images, D k Let be the k-th scene image, and s be the total number of scene images.

[0028] S6. Perform scene recognition on the scene image generated in step S5, and determine whether it matches the scene category information. If not, remove it. If it does, calculate the matching score between the scene image and the training semantic features through semantic relevance analysis and sort them. Keep the scene image with the highest matching score as the target scene image.

[0029] Scene recognition in this embodiment can be achieved through the following methods: A scene recognition model is constructed, sample images are acquired, and background sample images are obtained by performing background recognition on the sample images. The background sample images are then input into the scene recognition model to obtain a first recognition result. The training sample images are then input into the scene recognition model to obtain a second recognition result. The target model loss value is obtained based on the difference between the first and second recognition results. The model parameters of the scene recognition model are adjusted based on the target model loss value to obtain a trained scene recognition model. The scene image is then input into the trained scene recognition model to obtain the scene recognition result.

[0030] Scene recognition in this embodiment can also be achieved through the following other method: The semantic information of the scene image is obtained using a question-answering model; based on the obtained semantic information, the scene recognition result is obtained using a text recognition algorithm.

[0031] In this embodiment, semantic relevance analysis is implemented through the following method: The question-answering model is used to parse the scene image to obtain a set of scene semantic features D. kr , , where D kr For the k-th scene image D k A set of parsed scene semantic features. For the Dth kr For the k-th scene image D k The j-th scene semantic feature is parsed, and r is the total number of scene semantic features parsed; The obtained set of scene semantic features Dkr Perform feature mapping with a corresponding set of trained semantic features P; Calculate the feature similarity value between the scene semantic features and the training semantic features, and use the feature similarity value as the matching score. The feature similarity value is calculated using the following formula: , where W(D) k P) represents the k-th scene image D. k The feature similarity value between the scene semantic features and the training semantic features P, T ji for With p i The distance weights between the two features, c(j,i) are... With p i The distance between two features. T ji The value depends on p i The specific assignment method is as follows: first determine p i The corresponding part of speech, such as a noun, is T. ji The value is assigned as 1, and T is assigned if it is a verb. ji The value is assigned to 0.5; otherwise, T is assigned. ji The value is 0; then check p. i Does it belong to a scene category? If so, then T ji Multiply the value by 2, otherwise T ji The value remains unchanged.

[0032] S7. Display the target corpus data and the corresponding target scene image through the display device. The trainee performs oral training. The trainee's voice data is obtained through the sound receiving device. The voice data is analyzed to obtain the oral training score. If the oral training score is greater than the preset score threshold, the next target corpus data and the corresponding target scene image are displayed through the display device. If the oral training score is less than the preset score threshold, the trainee is prompted to retrain.

[0033] S8. Collect pronunciation error information from the trainee's oral training process, obtain a correction training dataset based on the foreign language words corresponding to the pronunciation error information, extract a set of matching target corpus data from the corpus sub-database based on the correction training dataset, and execute step S5. In this embodiment, the pronunciation error information is omission, mispronunciation, or inaccurate tone.

[0034] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for generating foreign language training content based on AIGC, characterized in that, Includes the following steps: S1. Construct a corpus database and a raw material library, wherein the corpus database contains corpus data for foreign language training, and the raw material library contains raw material data carrying text information and image information; S2. For the corpus database, classify the corpus data according to the pre-defined scene categories, and divide the corpus database into multiple corpus sub-databases based on the scene categories; S3. For the original material library, a material feature database is obtained using a question-answering model. The material feature database includes at least feature IDs, semantic features, and corresponding image regions. Based on the semantic features of the material feature database, scene classification is performed according to a pre-defined scene category, and the material feature database is divided into multiple material feature sub-databases based on the scene category. S4. Obtain the scenario category information for the foreign language training to be conducted, match the corpus sub-database and material feature sub-database based on the scenario category information, and randomly extract a set of target corpus data for the currently matched corpus sub-database. S5. For each target corpus data, use a question-answering model to parse it and obtain a corresponding set of training semantic features P. Where P is the set of training semantic features, p i Let m be the i-th training semantic feature, and m be the total number of training semantic features. Based on the trained semantic features, the system searches for matching semantic features and corresponding image region material Q in the corresponding material feature sub-database. Where Q is the set of training semantic features, Let j be the image region material, and n be the total number of image region materials. A set of scene images D is generated based on the AIGC content generation method. Where D is the set of scene images, D k Let s be the k-th scene image, and s be the total number of scene images; S6. Perform scene recognition on the scene image generated in step S5 to determine whether it matches the scene category information. If not, discard it; if so, calculate the matching score between the scene image and the trained semantic features through semantic relevance analysis and sort them. Keep the scene image with the highest matching score as the target scene image. The semantic relevance analysis is implemented through the following method: The question-answering model is used to parse the scene image to obtain a set of scene semantic features D. kr , , where D kr For the k-th scene image D k A set of parsed scene semantic features. For the Dth kr For the k-th scene image D k The j-th scene semantic feature is parsed, and r is the total number of scene semantic features parsed. The obtained set of scene semantic features D kr Perform feature mapping with a corresponding set of trained semantic features P; Calculate the feature similarity value between the scene semantic features and the training semantic features, and use the feature similarity value as the matching score. The feature similarity value is calculated using the following formula: Among them, W(D) k P) represents the k-th scene image D. k The feature similarity value between the scene semantic features and the training semantic features P, T ji for With p i The distance weights between the two features, c(j,i) are... With p i The distance between two features; S7. Display the target corpus data and the corresponding target scene image through the display device. The trainee performs oral training. The trainee's voice data is obtained through the sound receiving device. The voice data is analyzed to obtain the oral training score. If the oral training score is greater than the preset score threshold, the next target corpus data and the corresponding target scene image are displayed through the display device. If the oral training score is less than the preset score threshold, the trainee is prompted to retrain. S8. Collect pronunciation error information of trainees during oral training, obtain a correction training dataset based on the foreign language words corresponding to the pronunciation error information, extract a set of matching target corpus data from the corpus sub-database based on the correction training dataset, and execute step S5.

2. The method for generating foreign language training content based on AIGC as described in claim 1, characterized in that, The scene recognition in step S6 is achieved through the following method: A scene recognition model is constructed, sample images are obtained, background sample images are obtained by performing background recognition on the sample images, the background sample images are input into the scene recognition model to obtain a first recognition result, the training sample images are input into the scene recognition model to obtain a second recognition result, the target model loss value is obtained based on the difference between the first recognition result and the second recognition result, and the model parameters of the scene recognition model are adjusted based on the target model loss value to obtain a trained scene recognition model. The scene image is input into the trained scene recognition model to obtain the scene recognition result.

3. The method for generating foreign language training content based on AIGC as described in claim 2, characterized in that, The scene recognition in step S6 is achieved through the following method: Semantic information of scene images is obtained using a question-answering model; Based on the obtained semantic information, a text recognition algorithm is used to obtain scene recognition results.

4. A method for generating foreign language training content based on AIGC as described in claim 2 or 3, characterized in that, The acquisition of the material feature database using the question-answering model in step S3 is achieved through the following method: Construct a question-answering model, which includes an autocorrelation attention module and a cross-modal interaction attention module; Autocorrelation attention modules are used to extract features from both text and image information. Autocorrelation learning captures features between image regions and between text characters to assess semantic autocorrelation and update features. Given text information, word embedding encoding is used for representation, and word vector features are learned and extracted. For each given text information, the text feature Y is obtained, represented as: , where t j Let be the feature vector of the j-th word, and m be the total number of region features; Given image information T, a feature extraction network is trained to obtain image region features X, which are represented as follows: , where r i Let n be the i-th regional feature, and n be the total number of regional features. The semantic associations between images and text are learned and features are updated using a cross-modal interactive attention module; A cascaded approach is used to stack attention layers, and the image region features and text features are updated at multiple levels to obtain the final image region features and text features with clear semantic representation.

5. The method for generating foreign language training content based on AIGC as described in claim 4, characterized in that: The text information feature extraction adopts the gated recurrent unit method, and the image information feature extraction utilizes the Faster R-CNN network.

6. The method for generating foreign language training content based on AIGC as described in claim 5, characterized in that: The scenario category refers to the environment or situation in which the language expression takes place.

7. The method for generating foreign language training content based on AIGC as described in claim 6, characterized in that, In step S2, the corpus data is classified into scenarios according to pre-defined scenario categories, specifically as follows: Build a category rule library based on predefined scene categories; The pre-labeled corpus data samples are input into the scene classification model to obtain the trained scene classification model; The corpus data is input into the trained scene classification model to obtain category labels, and the corpus data is then labeled. Based on the category labels of the corpus data, the corresponding scene category is matched in the category rule base.

8. The method for generating foreign language training content based on AIGC as described in claim 1, characterized in that, The pronunciation error messages are omissions, mispronunciations, or inaccurate tones.

Citation Information

Patent Citations

  • Traditional cultural material library construction method and system based on artificial intelligence

    CN110990563A

  • Generative teaching resource system in virtual teaching scene and working method thereof

    CN117055724A