Portable Chinese learning machine resource retrieval method
Through the combination of natural language processing and deep learning, reinforcement learning, and collaborative filtering algorithms, the accuracy and personalized sorting of portable Chinese learning machine resource retrieval is achieved, solving the problem of low resource retrieval efficiency of traditional learning machine resource retrieval, and improving user learning experience and resource utilization efficiency.
Patent Information
- Application Number
- CN202510210131.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In resource retrieval, traditional portable Chinese learning machines have problems such as inaccurate instruction analysis, low resource matching efficiency, and lack of personalized sorting, which makes it difficult for users to quickly find learning resources that meet their needs.
The natural language processing toolkit is used for instruction analysis, combined with deep learning, reinforcement learning and collaborative filtering algorithms, and receive search instructions through voice, keyboard or handwriting input, perform multi-dimensional resource matching and personalized sorting, and use the inverted index structure to quickly find resources, combine user multi-dimensional operation data and similar user behavior to achieve personalized resource recommendations.
It improves the accuracy and efficiency of resource matching, can dynamically adjust the sorting according to changes in user interests, provides personalized resource recommendations, and significantly improves users' learning efficiency and experience.
Smart Images

Figure CN120336345A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of Chinese learning machines, and particularly relates to a resource retrieval method for a portable Chinese learning machine. Background Art
[0002] In the context of the current era of globalization, Chinese, as one of the languages with the largest number of speakers in the world, has become increasingly important, and the demand for learning Chinese has been growing globally. Portable Chinese learning machines, with their convenience and functionality, have become a popular choice among many Chinese learners, especially beginners, overseas Chinese learners, and those who need to utilize fragmented time for learning.
[0003] Traditional portable Chinese learning machines have significant drawbacks in resource retrieval. For example, in terms of instruction parsing, they mostly rely on simple natural language processing technology, which can only perform basic lexical and syntactic analysis and is difficult to understand the true semantics of user instructions; in resource matching, they use single-keyword matching and mainly search for resource names, ignoring in-depth content mining. Facing a large number of complex resources, the matching efficiency is poor, and it is easy to miss resources that match the content but do not contain the keyword in the name. There is also a lack of targeted strategies for different types of resources such as text, audio, and video. When screening and sorting, they often follow fixed rules, such as sorting according to resource popularity or upload time, without considering the user's personalized needs, interest preferences, and learning progress, which requires users to spend a lot of time screening out applicable resources and reduces learning efficiency.
[0004] Accordingly, this application proposes a resource retrieval method for a portable Chinese learning machine. Summary of the Invention
[0005] The purpose of the present invention is to solve the deficiencies existing in the prior art, and a resource retrieval method for a portable Chinese learning machine is proposed.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions:
[0007] A resource retrieval method for a portable Chinese learning machine includes the following steps:
[0008] S1. The user inputs a retrieval instruction, receives the retrieval instruction input by the user through voice, keyboard input, handwriting, etc., and uses technologies such as speech recognition and character recognition to convert it into processable text information;
[0009] S2. Instruction parsing, call the natural language processing toolkit, extract the core retrieval words through lexical analysis, syntactic analysis, and semantic analysis, and convert them into a unified retrieval format within the system;
[0010] S3. Repository matching: The intelligent calligraphy learning machine has teaching resource repositories for primary school unified textbooks, junior high school unified textbooks, Xiling Calligraphy Classroom, and brush calligraphy teaching to consolidate classroom knowledge, as well as rich resources such as stele inscription appreciation, calligraphy classes, idiom collections, and poetry appreciation. Based on the inverted index data structure, it can quickly search and retrieve resources in the repository that match the search terms, perform multi-dimensional matching, not only search for resource names but also delve into the resource content, or search for all relevant resources in the repository according to instructions and obtain the basic metadata of the resources;
[0011] S4. Screening and sorting: Collect multi-dimensional operation data of users and use deep learning algorithms, reinforcement learning algorithms, and collaborative filtering algorithms. Among them, the collaborative filtering algorithm collects users' learning behavior data to construct a user-resource matrix, calculates the similarity between users using the cosine similarity algorithm, finds similar user groups, analyzes the resources they use, and sorts them according to resource popularity and relevance to the current user;
[0012] S5. Result presentation: Display resources on the learning machine display screen in the form of lists, cards, pictures and texts.
[0013] Preferably, in the S1 step, a large number of speech sample datasets with different accents, speaking speeds, and contexts are used for training in the speech recognition technology.
[0014] Preferably, in the S2 step, the knowledge graph technology is used to assist semantic analysis, associate the extracted core search terms with relevant concepts in the knowledge graph, and mine more potential semantic information.
[0015] Preferably, in the step S3, the primary school unified textbooks include text reading, single-character learning, and calligraphy demonstrations;
[0016] The junior high school unified textbooks are upgraded to two calligraphy styles of regular script and running script, as well as the detailed mode and video mode;
[0017] The Xiling Calligraphy Classroom includes calligraphy teaching courses for eight semesters from grades three to six that are synchronized with the textbooks;
[0018] The brush calligraphy teaching covers the joint teaching of the five major calligraphy styles of seal script, official script, cursive script, running script, and regular script, as well as the four major schools of Ou, Yan, Liu, and Zhao;
[0019] The stele inscription appreciation collection includes various stele inscriptions from the pre-Qin period to the modern era;
[0020] The calligraphy classroom provides systematic teaching materials for hard-tipped pens and brush pens to help improve calligraphy skills;
[0021] The idiom collection and poetry appreciation include a rich collection of words and Tang poems and Song ci from primary school to high school.
[0022] Preferably, in the step S4, the deep learning algorithm uses a long short-term memory network to construct a user interest model to capture the long-term and short-term changes in user interests;
[0023] The reinforcement learning algorithm dynamically adjusts the weights of resource sorting by using the deep Q-network and policy gradient algorithms to adapt to the personalized needs of different users.
[0024] Preferably, in the step S5, the learning machine supports multi-language display, and translates information such as the name, introduction, and operation prompts of the resources into the corresponding language according to the language set by the user.
[0025] The present invention has the following beneficial effects:
[0026] 1. In the instruction parsing stage, with the help of natural language processing toolkits, lexical, syntactic, and semantic analyses are carried out, and various technologies such as pre-trained word vector models are combined to deeply mine potential semantics. By associating relevant concepts in the knowledge graph, the core retrieval terms are accurately extracted and converted into a unified format, improving the accuracy of understanding user intentions and guiding the direction for resource matching.
[0027] 2. The retrieval system is based on an inverted index structure and can quickly locate and match resources in the resource library. For different types of resources such as text, audio, and video, indexing strategies based on keywords, audio features, key-frame image features, and content descriptions are respectively adopted to carry out multi-dimensional matching and in-depth content search, improving the matching efficiency and accuracy and preventing resource omission.
[0028] 3. By comprehensively applying deep learning, reinforcement learning, and collaborative filtering algorithms, combined with multi-dimensional user operations, historical retrievals, resource feedback data, and similar user behaviors, personalized screening and sorting of retrieval results are realized. Deep learning captures changes in user interests to build a model, reinforcement learning adjusts weights according to feedback, and collaborative filtering analyzes resources of similar users, giving priority to displaying what is needed and improving the efficiency of obtaining high-quality resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 is the overall flowchart of a resource retrieval method for a portable Chinese learning machine proposed by the present invention;
[0030] Figure 2 is a partial code diagram based on deep learning in the first embodiment of the present invention;
[0031] Figure 3 is a partial code diagram based on reinforcement learning in the second embodiment of the present invention;
[0032] Figure 4 is a partial code diagram based on the collaborative filtering algorithm in the second embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0034] A resource retrieval method for a portable Chinese learning machine includes the following steps:
[0035] S1. The user inputs a retrieval instruction, receives the retrieval instruction input by the user through voice, keyboard input, handwriting, etc., and uses technologies such as speech recognition and character recognition to convert it into processable text information. Among them, the speech recognition technology is trained with a large number of speech sample data sets with different accents, speech rates, and contexts;
[0036] S2. Instruction parsing, call the natural language processing toolkit, extract the core retrieval terms through lexical analysis, syntactic analysis, and semantic analysis, convert them into a unified retrieval format within the system, and use the knowledge graph technology to assist semantic analysis, associate the extracted core retrieval terms with relevant concepts in the knowledge graph, and mine more potential semantic information;
[0037] S3. Resource library matching, the intelligent calligraphy learning machine has primary school unified compilation, junior high school unified compilation, Xiling Classroom, and writing brush teaching resource libraries to consolidate classroom knowledge, as well as rich resources such as stele inscription appreciation, calligraphy classroom, idiom encyclopedia, and poetry appreciation. Based on the inverted index data structure, it can quickly find resources matching the retrieval terms in the resource library, perform multi-dimensional matching, not only search for resource names, but also search deeply into the resource content, or find all relevant resources in the resource library according to the instruction, and obtain the basic metadata of the resources;
[0038] The primary school unified compilation includes text reading, single-character learning, and calligraphy demonstration. It starts from the basic sitting posture and pen-holding guidance, explains the new characters, and also designs functions such as following writing and dictation, step by step from recognition to writing;
[0039] The junior high school unified compilation is upgraded to two calligraphy styles of regular script and running script, as well as the detail mode and video mode, for word explanation, detail explanation, and character-by-character explanation.
[0040] The Xiling Classroom includes calligraphy teaching courses for the synchronous textbooks of eight semesters from the third to the sixth grades. It has complete content such as e-textbooks, teaching material courseware, and online micro-lesson teaching plans, making learning more efficient;
[0041] The writing brush teaching covers the joint teaching of the five calligraphy styles of seal script, official script, cursive script, running script, and regular script, and the four major schools of Ou, Yan, Liu, and Zhao. The course arrangement is carried out according to the basic strokes, radical components, shoulder frame structure, original post copying, thirty-six methods, and layout of the composition, with about 500 courses, and the famous artist demonstration courses are arranged step by step;
[0042] The rubbings appreciation section includes various rubbings from the pre-Qin period to the modern era, with approximately more than 4,000 single-character pictures reaching over 3.4 million. It can be searched in detail by script style, dynasty, and author, and with a vast amount of rubbings resources and high-definition pictures, detailed annotation and learning reference are both available;
[0043] The calligraphy classroom provides systematic teaching materials for both hard-tipped pens and writing brushes to assist in calligraphy improvement.
[0044] The idiom dictionary and poetry appreciation section includes a rich collection of words as well as Tang and Song poems from primary school to high school. At the same time, content such as The Book of Songs, Songs of Chu, and Yuefu are used as supplements, which can broaden the cultural perspective, enrich writing materials, and enhance cultural literacy;
[0045] S4. Screening and sorting, collecting multi-dimensional operation data of users, using deep learning algorithms, and constructing a user interest model using long short-term memory networks to capture the long-term and short-term changes of user interests;
[0046] The reinforcement learning algorithm uses deep Q networks and policy gradient algorithms to dynamically adjust the weights of resource sorting to meet the personalized needs of different users;
[0047] The collaborative filtering algorithm constructs a user-resource matrix by collecting user learning behavior data, calculates the similarity between users using the cosine similarity algorithm, finds similar user groups, analyzes the resources they use, and sorts them according to resource popularity and relevance to the current user;
[0048] S5. Result presentation, displaying resources in the form of lists, cards, pictures and texts on the learning machine display screen. At the same time, the learning machine also supports multi-language display. According to the language set by the user, information such as the name, introduction, and operation tips of the resources will be translated into the corresponding language.
[0049] Example 1:
[0050] Step 1: The user inputs a retrieval instruction
[0051] The user long-presses the voice input button on the learning machine and clearly says "Find learning materials and learning videos about calligraphy". The highly sensitive microphone built into the learning machine receives the voice signal and converts the analog voice signal into a digital signal through audio sampling and quantization techniques. The speech recognition module uses an end-to-end speech recognition model based on deep learning, an architecture that combines a convolutional neural network (CNN) and a recurrent neural network (RNN), to extract features and recognize the digital voice signal, and finally converts it into the text information "Find learning materials and learning videos about calligraphy".
[0052] Step 2: Instruction parsing
[0053] The learning machine processor calls the natural language processing (NLP) toolkit, first performs lexical analysis, and uses the word segmentation algorithm to split the sentence "Search for learning materials and learning videos about calligraphy" into independent words such as "Search", "About", "Calligraphy", "Learning Materials", and "Learning Videos". Then it performs syntactic analysis, builds a syntax tree, and clarifies the grammatical relationship between words, such as "Calligraphy" modifies "Learning Materials" and "Learning Videos". Finally, it performs semantic analysis through the pre-trained word vector model, maps the words to the vector space, understands the semantics, extracts the core search terms "Calligraphy", "Learning Materials", and "Learning Videos", and converts them into a unified search format and specific encoding form within the system.
[0054] Step 3: Resource library matching
[0055] The retrieval system is based on the inverted index data structure, which allows for quick searches in the resource library. For keywords such as "calligraphy", "learning materials", and "learning videos", the system locates resources including documents on calligraphy theory, texts appreciating calligraphy works, and calligraphy teaching videos, and also obtains basic metadata of the resources, such as file name, size, creation time, and video duration.
[0056] Step 4: Filter and sort
[0057] The deep learning algorithm module regularly collects the user's operation data on the learning machine, such as the history of browsing calligraphy-related content, the length of stay, the collection and like behavior of different calligraphy font learning materials, etc. These multi-dimensional data are preprocessed and converted into feature vectors suitable for model input. For example, the browsing time is normalized to the range of 0-1, and the collection behavior is represented by 0 and 1. The user interest model is constructed using the long short-term memory network (LSTM), and its basic formula is:
[0058] it=σ(W ii x t +b ii +W hi h t-1 +b hi )
[0059] ft=σ(W if x t +b if +W hf h t-1 +b hf )
[0060] ot=σ(W io x t +b io +W ho h t-1 +b ho )
[0061] ct = f t Θc t-1 + i t Θtanh(W ic x t + b ic + W hc h t-1 + b hc )
[0062] h t = o t Θtanh(c t )
[0063] Among them, x t is the current input, h t-1 is the hidden state at the previous moment, i t , f t , o t are the input gate, forget gate, and output gate respectively, c t is the memory cell, σ is the sigmoid function, Θ represents element-wise multiplication, and W and b are the weights and biases of the model.
[0064] By continuously training the model, it can accurately capture the user's interest preferences for calligraphy-related resources. According to the constructed interest model, the matched calligraphy-related resources are evaluated. Feature extraction and analysis are performed on the text content, audio descriptions, video tags, etc. in the resources, and the similarity score between them and the user interest model is calculated. Here, the cosine similarity formula is used to calculate the similarity:
[0065]
[0066] Among them, A and B are the resource feature vector and the user interest model vector respectively. Resources with high similarity scores, such as high-quality explanation videos that elaborate on the artistic conception of calligraphy and classic calligraphy texts with annotations of rare Chinese characters, are sorted from high to low according to the scores and preferentially displayed.
[0067] Step Five: Result Presentation
[0068] On the high-definition display screen of the learning machine, the resources are displayed in a combination of list and card forms. For calligraphy learning videos, the video cover image, video duration, play volume, brief introduction (the calligraphy font mainly explained in the video, the teaching teacher, etc.), and play button are displayed. For calligraphy learning materials, the material name, author, main content included (calligraphy history, font characteristics, etc.), and a preview of some content are displayed, and clicking can view the full text or download. At the same time, operation tips on how to play the video, download the material, etc. are provided on the interface.
[0069] Example Two:
[0070] Step 1: User enters a retrieval instruction
[0071] The user enters "calligraphy learning materials and videos" character by character by clicking the keys on the virtual keyboard of the learning machine. The learning machine receives the characters entered by the user in real time, performs real-time verification and error correction, and automatically prompts possible spelling mistakes and correction suggestions.
[0072] Step 2: Instruction parsing
[0073] The learning machine uses regular expressions and rule matching methods to perform a preliminary analysis of the input instruction, identifying key concepts such as "calligraphy", "learning materials", and "videos". Further, through a semantic understanding model, it determines that the user's core need is to obtain calligraphy-related learning materials and video resources, and converts them into a retrieval instruction format that can be understood by the system, including keywords and their weight information.
[0074] Step 3: Resource library matching
[0075] The retrieval system performs multi-dimensional matching in the resource library according to the parsed retrieval instruction. It not only searches for files whose resource names contain "calligraphy learning materials" and "calligraphy learning videos", but also delves into the resource content and uses a text search algorithm to find learning materials and teaching video resources containing calligraphy knowledge.
[0076] The matched resources include electronic documents of calligraphy theory books, PDF materials for calligraphy practice in different fonts, videos of famous calligraphers' demonstration writings, etc.
[0077] Step 4: Screening and sorting
[0078] The reinforcement learning algorithm maintains a state space, and the state s t ∈S includes the user's current retrieval instruction, historical retrieval records, and feedback data on different types of calligraphy resources (number of clicks, learning duration, marking and note-taking behaviors on the materials, etc.);
[0079] At the same time, a reward function R(s t ,a t ) is defined. If the user clicks and completely learns a calligraphy learning video and takes notes, a reward is given; if the user does not finish learning or has no interaction behavior after clicking, a reward R = -0.5 is given.
[0080] The algorithm selects an action a t according to the current state, that is, adjusts the weights of vocabulary matching degree, resource popularity, and user's historical usage behavior in the relevance algorithm. Here, the Q-learning algorithm is used, and its core formula is:[[]]
[0081]
[0082] Among them, a is the learning rate, and γ is the discount factor. According to the adjusted weights, the relevance score of each matching resource is recalculated, and the calligraphy learning resources of the multiple-choice question type with high scores are ranked at the top for display. The formula for calculating the relevance score is:
[0083] S core = w1 × MatchDegree + w2 × Popularity + w3 × UserHistory
[0084] Among them, w1, w2, and w3 are the adjusted weights, MatchDegree is the lexical matching degree, Popularity is the resource popularity, and UserHistory is the user's historical usage behavior score.
[0085] Step Five: Result Presentation
[0086] Display various calligraphy learning resources in a list form on the learning machine screen. Each resource entry displays the resource name, resource type ("Calligraphy Regular Script Learning Materials", "Calligraphy Running Script Teaching Videos"), difficulty level (represented by stars, 1 - 5 stars, 1 star for basic entry, 5 stars for advanced progression), and estimated learning time. For online learning video links, the words "Online Video" are specially marked, and the number of viewers and average evaluation score are displayed, facilitating users to quickly understand and select resources suitable for themselves.
[0087] Example Three:
[0088] Step One: User Input Retrieval Instruction
[0089] The user writes "Find calligraphy learning materials and videos" on the handwriting input area of the learning machine with a finger or a stylus. The handwriting recognition module of the learning machine uses image recognition - based technology to sample and extract features from the handwriting trajectory, converting the handwriting image into a digital signal. The convolutional neural network and other models are used to recognize the handwriting digital signal and convert it into text, "Find calligraphy learning materials and videos".
[0090] Step Two: Instruction Parsing
[0091] The learning machine first performs part - of - speech tagging on the input instruction to determine that "Find" is a verb, and "calligraphy learning materials" and "videos" are noun phrases. Using semantic analysis tools, it understands the semantics of "calligraphy learning materials" and "videos" and matches them with the classification labels in the resource library to determine that the retrieval scope is the calligraphy category in the cultural and artistic learning resources. The parsed information is converted into a retrieval instruction executable by the system, including keyword and category - limiting information.
[0092] Step Three: Resource Library Matching
[0093] According to the instructions, the retrieval system searches for all calligraphy learning materials and video resources in the resource library. These resources include text materials of calligraphy basic tutorials, e-books of appreciation of ancient calligraphy works, and a series of teaching videos from basic strokes to complete works. At the same time, the relevant metadata of each resource is obtained, including the name of the material, the author, the video lecturer, the applicable learning stage, etc.
[0094] Step 4: Filter and sort
[0095] Through the collaborative filtering algorithm module, a large amount of user learning behavior data is collected to build a user resource matrix, where M ij Represents the operation behavior of user i on resource j (whether it has been learned, marked as favorite, etc.). The cosine similarity algorithm is used to calculate the similarity between users. The formula is:
[0096]
[0097] Among them, sim(i,k) is the similarity between user i and user k. Find similar user groups with similar learning behaviors (frequently learning cultural and artistic resources, especially interested in calligraphy) as the current user. Analyze other resources that similar user groups often use when learning calligraphy, such as calligraphy practice tool recommendations and auxiliary materials for calligraphy appreciation.
[0098] Integrate calligraphy learning materials and video resources, as well as the selected related auxiliary resources, sort them according to the popularity of the resources (number of learning times, number of collection times) and the relevance to the current user, and put popular calligraphy learning videos and learning materials highly rated by similar users at the front. The relevance sorting formula can be expressed as:
[0099] Rank=w4×Popuarity+w5×Similarity
[0100] Among them, w4 and w5 are weights, Popularity is the popularity of the resource, and Similarity is the similarity score with the current user.
[0101] Step 5: Results presentation
[0102] On the learning machine display, resources are displayed in a graphic and textual manner. For calligraphy learning videos, the video cover, video name, lecturer introduction and play button are displayed; for calligraphy learning materials, the material cover picture, material name, main content introduction and download entrance are displayed; for related auxiliary resources, the resource name, brief usage description and acquisition method are displayed. It is convenient for users to quickly select and use.
[0103] It should be noted that to ensure that the experimental results are not interfered by the hardware, all learning machines adopt the same hardware configuration. The processor is Qualcomm Snapdragon 865, equipped with Adreno 650 GPU, 8GB of running memory and 256GB of UFS 3.1 storage memory; at the same time, an IPS screen with a size of 12.7 inches and a resolution of 2560×1600 is selected, as well as a 7000mAh lithium battery. In addition, front and rear cameras, dual speakers and microphones are equipped, and a variety of sensors are integrated to achieve intelligent interaction.
[0104] Among them, Comparative Example 1 adopts a simple keyword matching retrieval algorithm. When matching resources, it simply compares texts in the resource library based on the keywords input by the user. This method is too simple and direct to deeply understand the complex needs of users. Therefore, the resource matching rate is only 75.63%, and it is difficult to accurately locate resources that meet complex needs. Comparative Example 2 adopts a retrieval sorting method with fixed weights, and pre-sets the weights of factors such as keyword matching degree and resource popularity to calculate resource relevance. However, due to the diverse characteristics of different user needs, fixed weights are difficult to adapt flexibly, which makes its resource matching rate 78.21%, showing average performance.
[0105] Example 1 uses deep learning to construct a user interest model. By analyzing the user's long-term multi-dimensional behavior data, browsing history, learning duration, collection records, etc., it can accurately capture the user's interest preferences, so that the resource matching rate reaches 92.35%.
[0106] Example 2 uses the reinforcement learning algorithm, which can continuously optimize the weights of factors such as vocabulary matching degree, resource popularity, and user's historical usage behavior in the relevance algorithm according to the user's feedback on the retrieval results, such as the user's click behavior and learning completion degree. Its resource matching rate reaches 90.12%. Example 3 adopts a collaborative filtering algorithm. By constructing a user-resource matrix, calculating the similarity between users, and recommending resources based on the behavior of similar users, while mining potential interest resources of users, a matching rate of 88.45% is achieved, as shown in Table 1:
[0107] Table 1: Comparison table of learning resource matching, retrieval time, relevant resources, and duplicate resources
[0108]
[0109] Specifically, the deep learning model and the inverted index structure in the first embodiment can quickly process voice input and locate and match resources, with the shortest average retrieval time of 0.56s. In the second embodiment, rule matching and multi-dimensional matching strategies are used in the process of instruction parsing and resource matching, and the time is 0.62s. In the third embodiment, image recognition and reasonable semantic analysis are used to determine the retrieval range, and the time is 0.71s. In contrast, Comparative Example 1 is 1.25s and Comparative Example 2 is 1.08s. This shows that the present invention can quickly enable users to obtain the required resources instantly in resource retrieval, greatly improving the convenience and fluency of use. Especially for users who need to obtain information in a timely manner, this efficient retrieval experience is crucial.
[0110] In the first embodiment, a long short-term memory network is used to construct a user interest model, and the cosine similarity between the computing resources and the model is used for sorting. The recommendation accuracy rate reaches 85.23%. The long short-term memory network can capture the long-term and short-term changes of user interests, making the recommendation highly fit the personalized needs. In the second embodiment, the Q-learning algorithm of reinforcement learning is adopted, and the resource sorting weight is dynamically adjusted according to user feedback. The accuracy rate is 83.45%, and the recommendation result can be continuously optimized based on the real-time behavior of users. In the third embodiment, the collaborative filtering algorithm is used to find similar user groups by calculating the cosine similarity between users to recommend resources, and the accuracy rate is as high as 88.67%, giving full play to the value of group behavior data. However, the accuracy rates of Comparative Example 1 and Comparative Example 2 are only 55.34% and 60.21% respectively. Their recommendation methods that simply rely on rules or statistical information have poor effects. The high accuracy rate of the embodiments enables users to obtain resources that better fit their own interests and needs, improving the resource utilization rate and the learning and use effects.
[0111] The deep learning model in the first embodiment combined with similarity sorting can effectively screen resources, and the lowest occurrence rate of duplicate resources is 3.21%. The Q-learning algorithm in the second embodiment can also avoid some duplicates during the process of dynamically adjusting the weight, and the occurrence rate is 4.56%. The collaborative filtering algorithm in the third embodiment integrates and sorts when considering the resources of similar users, and the occurrence rate is 5.12%. However, Comparative Example 1 and Comparative Example 2 are as high as 15.67% and 12.34% respectively. The lack of an effective duplicate removal mechanism and personalized sorting algorithm results in users possibly seeing a large number of duplicate resources, affecting the use experience. The low duplicate rate of the embodiments provides users with a more diverse resource selection, avoiding the waste of time and energy by users on duplicate resources, and enabling users to access more valuable information.
[0112] It should be noted that in the comparative examples, Example 1 adopts a simple keyword matching method. When faced with fuzzy search terms, it can only perform a simple literal comparison, which makes it difficult to understand the user's true intention, so the score for the ability to process fuzzy search terms is only 2. Although Example 2 is slightly better than Example 1, it still has a major defect. It adopts a fixed weight sorting method. When processing fuzzy search terms, it cannot flexibly adjust the relevance weight of resources according to fuzzy semantics, so the score is 2.8. Example 1 constructs a user interest model based on deep learning. The deep learning algorithm can learn a large amount of user behavior data, so as to accurately understand the user's potential needs in different situations, and the ability to process fuzzy search terms reaches 4.5. Example 2 uses a reinforcement learning algorithm to continuously optimize the search strategy based on user feedback on the search results, and reaches 4.0 in the processing of fuzzy search terms. Example 3 adopts a collaborative filtering algorithm to recommend resources based on similar user behaviors. When faced with fuzzy search terms, it will refer to the choices and usage of similar users when searching for similar fuzzy terms, so that the ability to process fuzzy search terms reaches 3.5. As shown in Table 2:
[0113] Table 2: Comparison of fuzzy search terms, complex search terms and integration of learning plans
[0114]
[0115] Specifically, Example 1 stands out with a score of 4.5. It adopts a multi-dimensional search strategy, comprehensively considering factors such as resource type, content depth, and relevance. No matter how complex the user's search request is, it can accurately locate the resources that best meet the needs and provide accurate and comprehensive search results. Example 2 scored 4 points and can flexibly handle various complex search situations and efficiently meet user needs. Example 3 scored 3.8 points and can also screen out valuable resources. In contrast, Example 1 and Example 2 scored only 2 points and 2.5 points, respectively. The search mechanism is single and cannot cope with complex requests;
[0116] In terms of integration with user learning plans, Example 1 scored 4 points. By collecting information such as user learning history and goals, the search resources are deeply associated with the learning plan, allowing users to obtain content that matches their learning progress and goals, which strongly supports learning. Example 2 scored 3.5 points and Example 3 scored 3.8 points. They can also combine search resources with learning plans to a certain extent to help users plan and advance their learning. Comparative Examples 1 and 2 scored only 1 point and 2 points, ignoring the user's learning plan, and the search results are out of touch with the plan.
[0117] Generally speaking, the advantages of the embodiments are obvious. Their powerful ability to handle complex retrievals can save users' time in screening information and improve retrieval efficiency. Their high degree of integration with learning plans can provide users with resources that fit their learning progress, helping users plan their learning more reasonably and enhance learning effects, and meeting users' needs in learning resource retrieval and utilization in all aspects.
[0118] In summary, the embodiments of the present invention have obvious advantages in all aspects such as instruction parsing, resource matching, and result sorting, providing users with more efficient, accurate, and personalized retrieval services, and significantly enhancing users' retrieval experience and resource utilization efficiency.
[0119] Furthermore, as Figure 2 shown, Embodiment 1 uses the LSTM model of deep learning for resource screening and sorting, imports relevant libraries of numpy, sklearn, and tensorflow.keras, and first simulates user operation data and resource features. By constructing a model containing an LSTM layer and a fully connected layer and compiling it using the adam optimizer and mean squared error loss function, it can better process sequence data. After adjusting the user operation data into a three-dimensional shape suitable for LSTM input for training, a user interest vector is predicted and generated. Then, the resource features are traversed, the cosine similarity score between them and the user interest vector is calculated, and finally the resources are sorted according to the score and the index is output. Its advantages are significant. LSTM has a powerful memory function and can effectively capture the changes in users' interests in the long term and short term. Through the learning of a large amount of data by the deep learning model, it can more accurately reflect users' interest preferences, so as to screen out resources that better meet their personalized needs for users and improve the accuracy and pertinence of resource recommendations.
[0120] Even further, as Figure 3 shown, Embodiment 2 adopts the Q-learning algorithm of reinforcement learning, simulates states, resources, and Q-tables, and defines the learning rate and discount factor. During 20 iterations, states and actions are randomly selected, and the Q-table is updated according to the reward function. Then, weights are defined to calculate the relevance scores of resources, complete the sorting of resources, and output the results. The outstanding advantage of this algorithm is that it can dynamically adjust the weights of resource sorting according to users' real-time feedback. This means that the system can continuously optimize the resource sorting method according to users' actual behaviors and feedback during use to adapt to the personalized needs of different users in different scenarios, making the recommended resources more in line with users' current actual needs, and greatly improving the flexibility and adaptability of resource screening and sorting.
[0121] Even further, as Figure 4As shown, Example 3 is based on the collaborative filtering algorithm. The numpy and sklearn libraries are imported to simulate the user-resource matrix and the current user. By calculating the cosine similarity, similar user groups are found, and the resources used by these similar users are analyzed. At the same time, the resource popularity is simulated, weights are defined to calculate the correlation score, and finally the resources are sorted and the indexes are output. The advantage of this algorithm is that it makes full use of the behavior information of other users. In resource recommendation, the usage habits and preferences of similar users have certain reference value. The collaborative filtering algorithm can dig out these potential information and provide more resource recommendations that meet the needs of the current user. By combining the resource popularity and the usage of similar users, multiple factors can be considered comprehensively, making the recommendation results more comprehensive and accurate, and effectively improving the probability that users can obtain high-quality resources.
[0122] Generally speaking, by implementing code to enable reinforcement learning to adapt to real-time feedback and adjust strategies, and collaborative filtering to optimize recommendations using group behavior, a multi-level and all-round resource recommendation mechanism is jointly constructed, providing users with more high-quality and demand-oriented resource recommendation services, and improving the efficiency and experience of users in obtaining resources.
[0123] The above is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution of the present invention and its inventive concept, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. A resource retrieval method for a portable Chinese learning machine, characterized in that, It includes the following steps: S1. The user inputs a retrieval instruction, receives the retrieval instruction input by the user through voice, keyboard input, handwriting, etc., and uses technologies such as speech recognition and character recognition to convert it into processable text information; S2. Instruction parsing, call the natural language processing toolkit, extract the core retrieval terms through lexical analysis, syntactic analysis, and semantic analysis, and convert them into a unified retrieval format within the system; S3. Resource library matching, the intelligent calligraphy learning machine has primary school unified compilation, junior high school unified compilation, Xiling Classroom, and brush calligraphy teaching resource libraries to consolidate classroom knowledge, as well as rich resources such as stele inscription appreciation, calligraphy classroom, idiom encyclopedia, and poetry appreciation. Based on the inverted index data structure, it can quickly find resources matching the retrieval terms in the resource library, perform multi-dimensional matching, not only search for resource names, but also search deeply into the resource content, or search for all relevant resources in the resource library according to the instruction, and obtain the basic metadata of the resources; S4. Screening and sorting, collect multi-dimensional operation data of the user, and use deep learning algorithms, reinforcement learning algorithms, and collaborative filtering algorithms. Among them, the collaborative filtering algorithm collects the user's learning behavior data to construct a user-resource matrix, uses the cosine similarity algorithm to calculate the similarity between users, finds similar user groups, analyzes the resources they use, and sorts them according to resource popularity and relevance to the current user; S5. Result presentation, display the resources on the learning machine display screen in the form of lists, cards, pictures and texts, etc.
2. The resource retrieval method of a portable Chinese learning machine according to claim 1, wherein In the S1 step, a large number of speech sample data sets with different accents, speech rates, and contexts are used for training in the speech recognition technology.
3. A resource retrieval method for a portable Chinese learning machine according to claim 1, characterized in that, In the S2 step, knowledge graph technology is used to assist semantic analysis, associate the extracted core retrieval terms with relevant concepts in the knowledge graph, and mine more potential semantic information.
4. A resource retrieval method for a portable Chinese learning machine according to claim 1, characterized in that In the step S3, the primary school unified compilation includes text reading, single-character learning, and calligraphy demonstration; The junior high school unified compilation is upgraded to two calligraphy styles of regular script and running script, as well as the detailed mode and video mode; The Xiling Classroom includes calligraphy teaching courses for the eight semesters from the third to the sixth grades that are synchronized with the textbooks; The brush calligraphy teaching covers the joint teaching of the five major calligraphy styles of seal script, official script, cursive script, running script, and regular script and the four major schools of Ou, Yan, Liu, and Zhao; The stele inscription appreciation includes various stele inscriptions from the pre-Qin period to the modern era; The calligraphy classroom provides systematic teaching materials for both hard-tipped pens and brush pens to help improve calligraphy skills; The idiom encyclopedia and poetry appreciation include a rich collection of words and Tang poems and Song ci from primary school to high school.
5. A resource retrieval method for a portable Chinese learning machine according to claim 1, characterized in that, In the step S4, the deep learning algorithm uses a long short-term memory network to construct a user interest model to capture the long-term and short-term changes of the user's interest; The reinforcement learning algorithm uses a deep Q network and a policy gradient algorithm to dynamically adjust the weights of resource sorting to meet the personalized needs of different users.
6. A resource retrieval method for a portable Chinese learning machine according to claim 1, characterized in that, In the step S5, the learning machine supports multi-language display, and according to the language set by the user, translates the information such as the name, introduction, and operation prompts of the resources into the corresponding language.