A method and server for generating micro-courses based on OCR recognition

Through the micro-course generation method based on OCR recognition, professional terms and operation steps in maintenance records are extracted and integrated to generate personalized micro-course resource packages, which solves the problem of long training data update cycle in the existing technology and improves training efficiency and effectiveness.

CN119580289BActive Publication Date: 2025-05-27BEIJING COLAYA TECH & SERVICE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510138467.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-27
Estimated Expiration
2045-02-08

AI Technical Summary

Technical Problem

The existing technology requires a lot of labor when compiling and updating rail transit equipment maintenance training materials, which leads to a long update cycle and is difficult to timely reflect the latest maintenance experience, affecting the quality and efficiency of maintenance work.

Method used

The micro-course generation method based on OCR recognition is adopted, and the image data of the maintenance document is collected and recorded, optical character recognition processing is performed, professional terms, fault codes and operation step information are extracted, and structured knowledge data is integrated into structured knowledge data. A personalized micro-course resource package is generated based on the knowledge level of the target user and the position information parameters.

Benefits of technology

The update efficiency of training materials has been improved, the generated micro-course resource packages meet actual needs, and the training effect has been improved, ensuring that maintenance personnel can obtain the latest maintenance knowledge and skills in a timely manner.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119580289B_ABST
    Figure CN119580289B_ABST
Patent Text Reader

Abstract

A method and server for generating micro-courses based on OCR recognition, which relates to the field of image data processing. The method includes: collecting image data of maintenance record documents, performing optical character recognition processing on the image data of maintenance record documents to obtain digital text data; performing natural language processing on the digital text data according to a preset bidirectional encoder representation model, and integrating professional term information, fault code information, and operation step information into structured knowledge data; obtaining the knowledge level parameter and job information parameter of the target user, performing difficulty grading on the structured knowledge data according to the knowledge level parameter to obtain difficulty-graded knowledge data; performing knowledge scope screening on the difficulty-graded knowledge data according to the job information parameter to obtain target knowledge data; selecting teaching materials matching the target knowledge data from the knowledge resource library, and generating a micro-course resource package according to the teaching materials. Implementing this method can improve the update efficiency of training materials.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image data processing, and particularly to a method and server for generating micro-courses based on OCR recognition. Background Art

[0002] With the rapid development of urban rail transit systems, equipment maintenance training plays an increasingly important role in ensuring operational safety and improving service quality. As a highly specialized field, rail transit equipment maintenance involves a large amount of complex technical knowledge and operating specifications, which poses higher requirements for the training of maintenance personnel. In today's rapid development of information technology, intelligent training methods are gradually becoming an important trend in the industry's development.

[0003] In related technologies, unified training materials and fixed courses can be adopted. The training department usually organizes centralized lectures regularly or provides standardized e-courseware for maintenance personnel to study. During the process of sorting out training materials, trainers need to manually screen and sort historical maintenance records and then compile training content based on experience. This method requires a large amount of human resources.

[0004] However, since the compilation and update of training content require a large amount of manual work, the update cycle of training materials is long, the update efficiency of training materials is low, and it is difficult to reflect the latest maintenance experience in a timely manner. As a result, maintenance personnel cannot obtain the latest maintenance knowledge and skills in a timely manner, which affects the quality and efficiency of maintenance work. Summary of the Invention

[0005] This application provides a method and server for generating micro-courses based on OCR recognition, which is used to improve the update efficiency of training materials.

[0006] In a first aspect, this application provides a method for generating micro-courses based on OCR recognition, which is applied to a server. The method includes: collecting image data of maintenance record documents, performing optical character recognition processing on the image data of the maintenance record documents to obtain digital text data; performing natural language processing on the digital text data according to a preset bidirectional encoder representation model, extracting professional term information, fault code information, and operation step information, and integrating the professional term information, the fault code information, and the operation step information into structured knowledge data; obtaining the knowledge level parameter and the position information parameter of the target user, performing difficulty grading on the structured knowledge data according to the knowledge level parameter to obtain difficulty-graded knowledge data; performing knowledge scope screening on the difficulty-graded knowledge data according to the position information parameter to obtain target knowledge data; selecting teaching materials matching the target knowledge data from a knowledge resource library, and generating a micro-course resource package according to the teaching materials.

[0007] By adopting the above technical solution, the image data of the maintenance record document is collected, processed into digital text data through OCR, and then natural language processing is performed using a preset bidirectional encoder representation model. Information such as professional terms, fault codes, and operation steps is integrated into structured knowledge data. Then, the target knowledge data is screened according to the knowledge level and position information parameters of the target user, and finally a micro-course resource package is generated. The data collection provides a basis for subsequent processing. The OCR processing realizes document digitization. The model processing improves the information integration effect. The parameter screening ensures the personalization of the training content. The generated micro-course resource package improves the efficiency of generating training materials, meets the actual needs, and improves the training effect.

[0008] Combined with some embodiments of the first aspect, in some embodiments, the step of performing natural language processing on the digital text data according to the preset bidirectional encoder representation model, extracting professional term information, fault code information, and operation step information, and integrating the professional term information, the fault code information, and the operation step information into structured knowledge data specifically includes: inputting the digital text data into the preset bidirectional encoder representation model to generate a text vector containing context semantic information; extracting professional term information, fault code information, and operation step information from the text vector; establishing a first mapping relationship between the fault code information and the professional term information, and a second mapping relationship between the fault code information and the operation step information, and determining fault knowledge association data based on the first mapping relationship and the second mapping relationship; classifying and sorting the fault knowledge association data, and establishing a hierarchical structure including fault types, fault characteristics, and solutions according to the fault occurrence frequency, maintenance difficulty, and equipment importance; connecting the hierarchical structure in series according to the sequence of fault handling in the solution to form structured knowledge data.

[0009] By adopting the above technical solution, the digital text data is input into the preset bidirectional encoder representation model to generate a text vector, key information such as professional terms is extracted, mapping relationships are established to determine fault knowledge association data, classified and sorted, and a hierarchical structure is established based on fault characteristics, etc., and finally connected in series into structured knowledge data. The vector generated by the model considers context semantics, improves the accuracy of information extraction, the mapping relationships make the knowledge associations tight, and the hierarchical structure makes the knowledge logic clear, which helps to improve the systematicness and coherence of the knowledge in the training materials, thereby improving the training quality and effect, and helping trainees better master maintenance knowledge.

[0010] In some embodiments in combination with some embodiments of the first aspect, the steps of obtaining the knowledge level parameter and the job information parameter of the target user, grading the structured knowledge data according to the knowledge level parameter to obtain the difficulty-graded knowledge data, and screening the knowledge scope in the difficulty-graded knowledge data according to the job information parameter to obtain the target knowledge data specifically include: reading the working years data, professional qualification certificate data, and historical training record data of the target user from the user profile database, calculating the knowledge level score according to a pre-designed scoring rule, and using the knowledge level score as the knowledge level parameter; reading the job code, job responsibility list, and skill requirement list of the target user from the user profile database, and generating a job information parameter including job characteristic identifiers; grading the structured knowledge data according to the knowledge level parameter, dividing the structured knowledge data into primary, intermediate, and advanced levels to obtain the difficulty-graded knowledge data; calculating the matching degree between the job characteristic identifiers in the job information parameter and the knowledge labels in the difficulty-graded knowledge data, sorting the knowledge content according to the matching degree score, and selecting the knowledge content with a matching degree exceeding a preset matching threshold as the target knowledge data.

[0011] By adopting the above technical solution, data such as the working years of the target user are obtained from the user profile database to calculate the knowledge level score, relevant job information is read to generate the job information parameter, and based on this, the structured knowledge data is graded by difficulty and the knowledge scope is screened to obtain the target knowledge data. Multi-dimensional data is obtained to accurately evaluate the user's knowledge level. The generated parameters are used for difficulty grading to make the content match the user's ability, and the scope screening ensures relevance to the job requirements. Accurate evaluation and personalized customization improve the pertinence and practicality of training, enabling users to efficiently obtain matching knowledge and enhance their work ability and efficiency.

[0012] In some embodiments in combination with some embodiments of the first aspect, after the step of selecting teaching materials matching the target knowledge data from the knowledge resource library and generating a micro-course resource package according to the teaching materials, the method further includes: constructing a fault scenario simulation practice unit based on the fault code information and the operation step information; setting a plurality of checkpoints in the fault scenario simulation practice unit, each checkpoint corresponding to a key operation in the operation step information; recording the operation behavior data of the target user at each checkpoint, and comparing the operation behavior data with the standard operation steps in real time; when detecting an incorrect operation of the target user, extracting the professional term information and operation step information related to the standard operation steps from the target knowledge data to generate a prompt content, and sending it to the target client.

[0013] By adopting the above technical solution, a fault scenario simulation exercise unit is constructed based on the fault code information and operation step information. Key operations corresponding to checkpoints are set therein, the operation behavior data of the target user is recorded and compared with the standard operation. When an incorrect operation is detected, relevant information is extracted from the target knowledge data to generate prompt content and sent to the client. The constructed simulation unit can reproduce the real maintenance scenario, the checkpoints ensure the standardization of operations, the recording and comparison can promptly detect errors, and extracting information from the target knowledge data to generate prompt content ensures the accuracy and pertinence of the prompts.

[0014] In combination with some embodiments of the first aspect, in some embodiments, after the step of extracting the professional term information and operation step information related to the standard operation step from the target knowledge data to generate prompt content when detecting an incorrect operation of the target user, the method further includes: obtaining the score situation of the target user in the fault scenario simulation exercise unit; adjusting the difficulty level and knowledge point distribution of the micro-course resource package according to the score situation.

[0015] By adopting the above technical solution, the score situation of the target user in the fault scenario simulation exercise unit is obtained, and the difficulty level and knowledge point distribution of the micro-course resource package are adjusted according to the score. The score situation intuitively reflects the user's learning effect and operation proficiency. Adjusting the difficulty level based on this can make the course always adapt to the user's ability, avoid the learning enthusiasm being affected by too high or too low difficulty, and the adjustment of the knowledge point distribution can focus on the user's weak links, strengthen key knowledge, make the course content more reasonable, and thus optimize the user's learning path.

[0016] In combination with some embodiments of the first aspect, in some embodiments, after the step of selecting teaching materials matching the target knowledge data from the knowledge resource library and generating a micro-course resource package according to the teaching materials, the method further includes: statistically analyzing the learning behavior data of multiple target users using the same micro-course resource package, and extracting the knowledge point access frequency, average learning duration, and repeated learning times in the learning behavior data; performing importance annotation on the professional term information and operation step information in the structured knowledge data according to the knowledge point access frequency, the average learning duration, and the repeated learning times to obtain annotation information; adjusting the presentation order and display duration of the teaching materials in the micro-course resource package based on the annotation information.

[0017] By adopting the above technical solution, the learning behavior data of multiple target users using the same micro-course resource package are statistically analyzed, and information such as the frequency of knowledge point visits is extracted. Based on this, the importance of professional terms and operation steps in the structured knowledge data is annotated, and the presentation order and display duration of teaching materials in the micro-course resource package are adjusted based on the annotations. Statistical analysis of multi-user learning behavior data can explore the actual value of knowledge, and importance annotation provides a basis for adjusting teaching materials. Reasonable adjustment of presentation order and display duration can highlight key knowledge and attract user attention.

[0018] In combination with some embodiments of the first aspect, in some embodiments, after the step of selecting teaching materials that match the target knowledge data from the knowledge resource library and generating a micro-course resource package based on the teaching materials, the method also includes: collecting problem feedback submitted by the target user during the learning process, and associating problem points that appear more than a preset frequency with the target knowledge content in the target knowledge data; and updating the content structure of the micro-course resource package based on the target knowledge content.

[0019] By adopting the above technical solutions, we can collect user feedback, screen high-frequency problem points and associate them with target knowledge content, and update the content structure of the micro-course resource package. Problem feedback reflects the learning difficulties of users, and the determination of high-frequency problem points focuses on key issues. The association with target knowledge content achieves precise optimization, making the course content more in line with user needs, strengthening weak links, and making knowledge explanation more in-depth and thorough. It improves the quality of training materials, makes users learn more smoothly, effectively improves learning effects, enhances the pertinence and practicality of training, and helps users better master knowledge.

[0020] In a second aspect, an embodiment of the present application provides a server, comprising: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the server to execute the method described in the first aspect and any possible implementation method of the first aspect.

[0021] In a third aspect, an embodiment of the present application provides a computer program product comprising instructions, which, when executed on a server, enables the server to execute the method described in the first aspect and any possible implementation of the first aspect.

[0022] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, comprising instructions, which, when executed on a server, enable the server to execute the method described in the first aspect and any possible implementation of the first aspect.

[0023] Understandably, the server provided in the second aspect above, the computer program product provided in the third aspect, and the computer storage medium provided in the fourth aspect are all used to execute the method provided in the embodiments of the present application. Therefore, the beneficial effects they can achieve can refer to the beneficial effects in the corresponding method, which will not be elaborated here.

[0024] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:

[0025] 1. In the present application, by collecting the image data of maintenance record documents, processing it into digital text data through OCR, and then performing natural language processing using a preset bidirectional encoder representation model, information such as professional terms, fault codes, and operation steps is integrated into structured knowledge data. Then, target knowledge data is screened according to the knowledge level and job information parameters of the target user, and finally a micro-course resource package is generated. The data collection provides a basis for subsequent processing, the OCR processing realizes document digitization, the model processing improves the information integration effect, the parameter screening ensures the personalization of training content, and the generated micro-course resource package improves the efficiency of generating training materials, fits the actual needs, and improves the training effect.

[0026] 2. In the present application, by constructing a fault scenario simulation practice unit based on fault code information and operation step information, setting checkpoints corresponding to key operations therein, recording the operation behavior data of the target user and comparing it with the standard operation, when an incorrect operation is detected, relevant information is extracted from the target knowledge data to generate prompt content and sent to the client. Constructing the simulation unit can reproduce the real maintenance scenario, the checkpoints ensure the standardization of operations, the recording and comparison can promptly detect errors, and extracting information from the target knowledge data to generate prompt content ensures the accuracy and pertinence of the prompts.

[0027] 3. In the present application, by obtaining the score situation of the target user in the fault scenario simulation practice unit, the difficulty level and knowledge point distribution of the micro-course resource package are adjusted according to the score. The score situation intuitively reflects the user's learning effect and operation proficiency. Adjusting the difficulty level based on this can make the course always adapt to the user's ability, avoiding the influence of too high or too low difficulty on learning enthusiasm. The adjustment of the knowledge point distribution can focus on the user's weak links, strengthen key knowledge, make the course content more reasonable, and thus optimize the user's learning path. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 is a flow schematic diagram of a micro-course generation method based on OCR recognition in the embodiments of the present application;

[0029] Figure 2 is another flow schematic diagram of a micro-course generation method based on OCR recognition in the embodiments of the present application;

[0030] Figure 3It is another process schematic diagram of the micro-course generation method based on OCR recognition in the embodiments of the present application;

[0031] Figure 4 It is a schematic structural diagram of an entity device of the server in the embodiments of the present application. Detailed implementation manners

[0032] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. As used in the specification of the present application, the singular forms "a", "an", "the above", "the" and "this" are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in the present application refers to any or all possible combinations including one or more of the listed items.

[0033] Hereinafter, the terms "first" and "second" are only used for descriptive purposes, and cannot be understood as implying or suggesting relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.

[0034] It should be noted that currently, in urban rail transit, equipment maintenance is a complex and highly specialized task, involving a large number of mechanical devices of different types and specifications. Traditional maintenance training often relies on paper manuals or fixed electronic tutorials, which are slow to update and difficult to adapt to the rapidly changing equipment technology and the growing personalized training needs. With the development of information technology, especially in the application of optical character recognition (OCR) technology and large machine learning models, it has become possible to automatically process and understand unstructured documents, providing technical support for the intelligent generation of personalized work order parsing and training courses.

[0035] For ease of understanding, the application scenarios of the embodiments of the present application are introduced below.

[0036] Currently, there are usually two ways to process the data of maintenance work orders: one is to manually input the data into the system, which is time-consuming and laborious and prone to errors; the other is to use a simple text matching tool to automatically extract key words, but this method has low accuracy and cannot deeply understand and parse the content of the work order. In addition, there are rule-based expert systems for generating specific types of maintenance guides, but this requires a large number of rules to be defined in advance, with high maintenance costs and poor flexibility.

[0037] Taking the training of a certain subway company as an example, the training personnel regularly collect paper maintenance records, manually screen valuable information, and compile training materials or produce electronic courseware based on experience. This method requires a large amount of manpower and time. For example, when updating the maintenance training materials for the signal system, the training personnel have to check a large number of paper records one by one, sort out the key information, and then compile it into a unified textbook. However, due to the cumbersome manual operation and the long data update cycle, often long after new maintenance experience appears, the training materials are still not updated. The maintenance personnel can only operate based on old knowledge and may not be able to handle new faults in a timely and accurate manner, affecting the normal operation of the subway.

[0038] Still taking this subway company as an example, after adopting the method of this application, the server automatically collects the image data of the maintenance record documents, such as obtaining the image of the signal system maintenance record through a scanning device. After OCR processing, it is converted into digital text data, and then the information such as professional terms, fault codes, and operation steps is extracted by using a bidirectional encoder representation model, and integrated into structured knowledge data. According to the knowledge level parameters of the maintenance personnel (such as working years, training experience, etc.) and the position information parameters (such as position code, skill requirements, etc.), the suitable target knowledge data is screened out to generate a micro-course resource package. In this way, the maintenance training materials for the signal system can be updated quickly, and the maintenance personnel can obtain the latest knowledge in a timely manner. For example, for the new signal fault handling method, they can quickly learn and apply it to actual work, improving the maintenance efficiency and ensuring the safety of subway operation.

[0039] For the convenience of understanding, the method provided in this implementation will be described in terms of process in combination with the above scenario. Please refer to Figure 1 , which is a schematic flow diagram of the micro-course generation method based on OCR recognition in the embodiment of this application.

[0040] S101. Collect the image data of the maintenance record documents, and perform optical character recognition processing on the image data of the maintenance record documents to obtain digital text data.

[0041] Among them, the maintenance record document refers to the written material that records the equipment maintenance process, including information such as maintenance time, maintenance object, fault phenomenon, and treatment method, and usually exists in the form of paper documents, photos, or scanned copies. The image data refers to the digital image file formed after digitizing the scan or photograph of these documents, and the common formats include JPG, PNG, etc. Optical character recognition (OCR) is a technology that converts the text in the image into editable text, and is realized through steps such as image preprocessing, character segmentation, feature extraction, and pattern recognition. The digital text data refers to the machine-readable text information obtained after OCR processing.

[0042] Perform this step when it is necessary to convert paper-based maintenance records into a processable digital form. Specifically, the server first collects images of maintenance record documents through devices such as scanners and cameras, or obtains existing document scans from a document management system. Then, preprocess the collected images, including operations such as image denoising, skew correction, and binarization, to improve the image quality. Next, call an OCR engine (such as Tesseract) to perform character recognition on the processed images, converting the text information in the images into an editable text format, including text content, location information, etc. Finally, post-process the recognition results, such as correcting spelling mistakes and standardizing the format, to obtain standardized digital text data.

[0043] In some embodiments, document digitization can be achieved through a deep learning-based OCR method: First, use a pre-trained document image classification model to identify the document type and adopt corresponding preprocessing strategies according to different types; then use a text detection model based on a convolutional neural network to locate the text regions in the document and generate text box coordinates; finally, use an OCR recognition model based on a recurrent neural network to recognize the text content line by line and output digital text. Optionally, a traditional rule-based OCR method can also be used: First, perform grayscale and binarization processing on the image, and use projection analysis methods for layout analysis and row-column segmentation; then use a template matching algorithm to recognize individual characters; finally, perform post-processing through dictionary matching and context analysis to improve the recognition accuracy. It can be understood that other OCR technologies can also be used to achieve document digitization, which is not limited here.

[0044] S102. Perform natural language processing on the digital text data according to a preset bidirectional encoder representation model, extract professional term information, fault code information, and operation step information, and integrate the professional term information, the fault code information, and the operation step information into structured knowledge data.

[0045] Among them, the bidirectional encoder representation model is a neural network architecture that can consider both the forward and backward context information of the text sequence. Representative models include BERT, RoBERTa, etc. Professional term information refers to professional vocabulary in the maintenance field, such as equipment names, part names, etc. Fault code information is standardized fault numbers and descriptions. Operation step information includes specific maintenance operation processes and methods. Structured knowledge data is a data set that organizes the extracted information according to a predefined format.

[0046] This step is executed after obtaining digital text data for extracting valuable maintenance knowledge from the text. Specifically, the server inputs the digital text obtained by OCR into a pre-trained bidirectional encoder model (such as BERT). The model first tokenizes the text, converting each word into a word vector. Then, through a multi-layer Transformer structure, it performs bidirectional encoding on the word vector sequence, considering the context relationship of each word in the entire sequence. Finally, it outputs a vector representation of the text, which contains the semantic information of the text. Then, various types of information are extracted from the vector through an information extraction module, and the information is integrated into a structured format.

[0047] In some embodiments, information extraction can be achieved through a rule-based method: First, establish a professional term dictionary, a fault code library, and an operation step template; then, identify relevant information from the text through regular expressions and keyword matching; finally, organize the extracted information into a structured form according to a predefined pattern. Optionally, it can also be achieved through deep learning methods: First, use a pre-trained named entity recognition model to identify professional terms; then, use a BERT-based sequence labeling model to identify fault codes and operation steps; finally, obtain the relationships between entities through a relation extraction model to construct a knowledge graph. It can be understood that other natural language processing technologies can also be used to achieve information extraction and structuring, which are not limited here.

[0048] S103. Obtain the knowledge level parameter and the position information parameter of the target user, and classify the structured knowledge data according to the knowledge level parameter to obtain the difficulty-level classified knowledge data.

[0049] Among them, the target user refers to an individual maintenance personnel who needs to receive training. The knowledge level parameter refers to an index used to quantitatively evaluate the user's mastery of professional knowledge, including information such as working years, certification status, and training experience. The position information parameter is used to represent the working position characteristics of the user, including data such as position code, responsibility scope, and skill requirements. The structured knowledge data refers to a set of systematically organized maintenance knowledge content. Difficulty classification refers to the process of dividing knowledge content into different levels according to complexity. The difficulty-level classified knowledge data represents a knowledge system organized according to difficulty levels.

[0050] This step is executed when personalized knowledge customization is required based on user characteristics after obtaining structured knowledge data. Specifically, the server first reads data such as the working years, professional qualification certificates, and historical training records of the target user from the user profile database. Then, according to the preset scoring rules, these data are converted into standardized knowledge level scores. At the same time, information such as the user's job code, job responsibility list, and skill requirements is obtained to generate job characteristic data. Next, the knowledge level scores are used to evaluate the difficulty of the structured knowledge data, and the knowledge content is divided into three levels: elementary, intermediate, and advanced, forming a knowledge system with difficulty grading.

[0051] In some embodiments, knowledge level evaluation and difficulty grading can be achieved in various ways: Optionally, a rule-based scoring method is adopted: First, the basic scores for working years, certificate levels, and training scores are set; then, the weighted total score is calculated according to the weights of each score; finally, the total score is mapped to the score intervals of different difficulty levels. Optionally, a machine learning method is adopted: First, a large amount of knowledge level data and performance evaluation data of users are collected; then, a classification model is trained to learn the correspondence between knowledge level and difficulty level; finally, the model is used to predict the knowledge level of new users and match the content difficulty. It can be understood that other evaluation methods can also be used to achieve knowledge difficulty grading, which is not limited here.

[0052] S104. Screen the knowledge scope in the knowledge data with difficulty grading according to the job information parameters to obtain the target knowledge data.

[0053] Among them, the job information parameters represent the job characteristics of the user, including dimensions such as job category, job responsibilities, and skill requirements. The knowledge data with difficulty grading refers to the set of knowledge content divided according to difficulty levels. Knowledge scope screening refers to the process of selecting relevant knowledge according to job requirements. The target knowledge data represents the set of knowledge content that matches the job after screening. The matching degree refers to the degree of relevance between the knowledge content and the job requirements.

[0054] This step is executed when the training content scope needs to be determined according to job characteristics after completing the knowledge difficulty grading. Specifically, the server first converts the job information parameters into standardized feature vectors, including job category codes, responsibility item codes, skill level codes, etc. Then, in the knowledge data with difficulty grading, the similarity between each knowledge unit and the job feature vector is calculated to generate a matching degree score. Next, according to the preset matching threshold, the knowledge content with a higher matching degree is screened to ensure that the selected knowledge is highly relevant to the job requirements. Finally, the screened knowledge content is organized into a target knowledge data set.

[0055] In some embodiments, the knowledge scope screening can be achieved in various ways: Optionally, a rule-based mapping method is adopted: First, establish a job-knowledge correspondence table and define the knowledge items required for each job; then query the corresponding knowledge list according to the user's job code; finally, perform cross-screening in combination with the difficulty level to obtain the final knowledge set. Optionally, a semantic matching method is adopted: First, convert the job description and knowledge content into semantic vectors; then calculate the cosine similarity between the vectors; finally, screen the matching knowledge content according to the similarity threshold. It can be understood that other screening methods can also be used to determine the knowledge scope, which is not limited here.

[0056] S105. Select teaching materials that match the target knowledge data from the knowledge resource library, and generate a micro-course resource package according to the teaching materials.

[0057] Among them, the knowledge resource library refers to a data system that stores various teaching materials, including multimedia resources such as text, pictures, videos, and animations. Teaching materials refer to specific teaching content carriers, such as knowledge point explanation documents, operation demonstration videos, exercise questions, etc. The target knowledge data is used to represent the set of knowledge content related to the job after screening. The matching degree refers to the degree of relevance between the teaching materials and the target knowledge. The micro-course resource package represents a complete learning unit, including components such as teaching objectives, learning content, practice tests, and evaluation plans. The teaching law refers to the teaching content organization method that follows the cognitive law.

[0058] After determining the target knowledge data, this step needs to be executed when constructing a specific training course. Specifically, the server first divides the target knowledge data into several learning units according to knowledge points. For each learning unit, retrieve relevant teaching materials from the knowledge resource library, including text explanations, graphic illustrations, video demonstrations, etc. The selected teaching materials need to meet the requirements of complete knowledge point coverage, diverse expression forms, and moderate difficulty levels. Organize the selected materials according to the teaching law to form a standard course structure including five links: "introduction - explanation - example - exercise - summary". The duration of each link is controlled within 5 - 10 minutes to ensure the learning effect. Finally, package all the resources to generate a complete micro-course resource package including course information, learning objectives, teaching content, and assessment criteria.

[0059] In some embodiments, teaching material matching and micro-course generation can be achieved in various ways: Optionally, a matching method based on vector similarity is adopted: First, convert the target knowledge content and teaching materials into feature vectors; then calculate the similarity scores between the vectors and select the material with the highest similarity; finally, organize the content according to the instructional design template to generate a standardized course package. Optionally, an association method based on a knowledge graph is adopted: First, construct a domain knowledge graph and establish the association relationship between knowledge points and teaching materials; then search for relevant materials according to the positions of the knowledge points in the graph; finally, design the course structure according to the learning path to form a systematic learning resource package. It can be understood that other methods can also be used to select teaching materials and generate courses, which are not limited here.

[0060] The following further describes the method provided in this embodiment in a more specific process. Please refer to Figure 2 , which is another process schematic diagram of the micro-course generation method based on OCR recognition in the embodiments of the present application.

[0061] S201. Collect the image data of the maintenance record document, and perform optical character recognition processing on the image data of the maintenance record document to obtain digital text data.

[0062] The maintenance record document refers to the written material recording the equipment maintenance process, including information such as maintenance time, maintenance object, fault phenomenon, treatment method, etc., and usually exists in the form of paper documents, photos or scanned copies. The image data refers to the digital image file formed after digitizing and scanning or photographing these documents, and common formats include JPG, PNG, etc. Optical character recognition (OCR) is a technology that converts the text in an image into editable text, and is achieved through steps such as image preprocessing, character segmentation, feature extraction, and pattern recognition. The digital text data refers to the machine-readable text information obtained after OCR processing.

[0063] In the specific execution process, the server first collects the image of the maintenance record document through devices such as scanners and cameras, or obtains the existing document scanned copy from the document management system. Then, preprocess the collected image, including operations such as image denoising, skew correction, and binarization, to improve the image quality. Then call the OCR engine (such as Tesseract) to perform character recognition on the processed image, and convert the text information in the image into an editable text format, including text content, position information, etc. Finally, post-process the recognition result, such as correcting typos and standardizing the format, to obtain standardized digital text data.

[0064] S202. Input the digital text data into a preset bidirectional encoder representation model to generate a text vector containing context semantic information.

[0065] The bidirectional encoder representation model is a neural network architecture that can consider both the forward and backward context information of a text sequence. Representative models include BERT, RoBERTa, etc. Context semantic information refers to the meaning of words in a specific context, including the dependency relationships and semantic associations between words. A text vector is a numerical representation of text that can capture the semantic features of the text, facilitating subsequent processing.

[0066] The server inputs the digital text obtained by OCR into a pre-trained bidirectional encoder model (such as BERT). The model first tokenizes the text, converting each word into a word vector. Then, through a multi-layer Transformer structure, it performs bidirectional encoding on the word vector sequence, considering the context relationship of each word in the entire sequence. Finally, it outputs a vector representation of the text. This vector contains the semantic information of the text and usually has a dimension of 768 or 1024. This vector representation preserves the semantic relationships in the text, which is beneficial for subsequent information extraction.

[0067] S203. Extract professional term information, fault code information, and operation step information from the text vector.

[0068] Professional term information refers to professional vocabulary in the maintenance field, such as equipment names, part names, etc. Fault code information is standardized fault numbers and descriptions. Operation step information includes specific maintenance operation processes and methods. All this information is contained in different dimensions of the text vector.

[0069] The server extracts information from the text vector in the following ways: uses a pre-trained named entity recognition model to identify professional terms, including equipment names, part names, tool names, etc.; identifies standard format fault codes through regular expressions or rule matching; uses a sequence annotation model to identify the temporal relationships of operation steps and extract the complete operation process. For each type of information, a corresponding standardized format is established to ensure the consistency and usability of the extraction results. The extracted information will be used for subsequent knowledge construction and course generation.

[0070] S204. Establish a first mapping relationship between the fault code information and the professional term information and a second mapping relationship between the fault code information and the operation step information, and determine fault knowledge association data based on the first mapping relationship and the second mapping relationship.

[0071] Fault code information refers to standardized fault numbers and descriptions, such as "E001 - Main motor overload". Professional term information includes standardized terms such as equipment names, part names, and professional concepts, such as "frequency converter", "bearing", etc. Operation step information is the specific operation guidance during the repair process, including operation actions, tool usage, parameter settings, etc. The mapping relationship refers to the corresponding connection between two types of information, established through association rules or semantic similarity. Fault knowledge association data refers to the knowledge network formed through the mapping relationship, containing the complete association information of fault codes, related terms, and processing steps.

[0072] In the specific implementation, the server first establishes the first mapping relationship between fault codes and professional terms: by analyzing the professional terms appearing in the fault code descriptions, establish the association between fault codes and related equipment, parts, parameters, etc. terms, forming a data structure of {fault code: [related term 1, related term 2, ...]}. Then establish the second mapping relationship between fault codes and operation steps: associate each fault code with the corresponding sequence of processing steps, forming a data structure of {fault code: [step 1, step 2, ...]}. Finally, integrate the two mapping relationships to form fault knowledge association data centered on fault codes, containing related terms and processing steps, with the data structure of {fault code: {related terms: [term 1, term 2, ...], operation steps: [step 1, step 2, ...]}}.

[0073] S205. Classify and organize the fault knowledge association data, and establish a hierarchical structure including fault types, fault characteristics, and solutions according to the fault occurrence frequency, repair difficulty, and equipment importance.

[0074] The fault occurrence frequency refers to the number of times a specific fault appears in the historical records, used to reflect the commonness of the fault. The repair difficulty is a comprehensive assessment based on factors such as operation complexity, required skill level, and time consumption. The equipment importance is determined based on the role of the equipment in the system, the scope of the fault impact, etc. The fault type is the classification of the nature of the fault, such as mechanical fault, electrical fault, etc. The fault characteristics include the description of the manifestation form, influencing factors, etc. of the fault. The solution is a complete repair plan including processing steps, required tools, precautions, etc. The hierarchical structure is a tree-like structure that organizes relevant information according to logical relationships.

[0075] The server classifies and organizes the fault knowledge association data: First, count the occurrence times of each fault and calculate the frequency; evaluate the maintenance difficulty score according to factors such as the number of operation steps and technical requirements; determine the equipment importance score in combination with the equipment function importance and influence range. Then classify the faults according to these three dimensions and construct a three-layer structure: The first layer is classified by fault type; the second layer contains the specific fault characteristics under each type; the third layer is the corresponding solution. The data structure is {fault type: {fault characteristics: {frequency: value, difficulty: value, importance: value, solution: content}}}.

[0076] S206. Connect the hierarchical structure in sequence according to the sequence of fault handling in the solution to form structured knowledge data.

[0077] The sequence of fault handling in the solution refers to the temporal dependence relationship of maintenance operations, including steps that must be executed in sequence and steps that can be executed in parallel. Connecting the hierarchical structure means reorganizing the scattered hierarchical nodes according to the operation sequence to form a directed knowledge graph. Structured knowledge data is a knowledge representation after normalization processing, including elements such as entities, relationships, and attributes.

[0078] The server serializes the solution part in the hierarchical structure: Analyze the dependence relationship of the operation steps in each solution, identify the sequence of steps that must be executed in sequence; add preconditions and subsequent step identifiers to each step; mark parallel operations as steps at the same level. Then reorganize the hierarchical structure according to the operation sequence to generate structured data containing a complete knowledge link. The data structure is {knowledge unit: {type: value, characteristics: value, pre: [], post: [], parallel: [], attributes: {frequency: value, difficulty: value, importance: value}, content: value}}. This structure not only retains the knowledge classification system but also reflects the sequence relationship of operations.

[0079] S207. Read the working years data, professional qualification certificate data, and historical training record data of the target user from the user profile database, calculate the knowledge level score according to the pre-designed scoring rules, and use this knowledge level score as the knowledge level parameter.

[0080] The user profile database is a structured data storage system that stores user basic information and career development information. The working years data records the working duration of the user in the relevant position in years. The professional qualification certificate data includes information such as the certificate name, level, and acquisition time, such as "Senior Maintenance Engineer Certificate". The historical training record data includes information such as the training course name, completion time, and assessment score. The pre-designed scoring rules are a quantitative standard for evaluating the user's knowledge level, including the weights and calculation methods of various indicators. The knowledge level score is a quantitative representation of the user's mastery of professional knowledge, usually using a percentage system.

[0081] The server reads the target user information from the user profile database through an SQL query. First, it extracts the work experience data and calculates the basic score based on 10 points per year. Then it analyzes the professional qualification certificate data, with 15 points for a junior certificate, 25 points for an intermediate certificate, and 40 points for a senior certificate. Next, it counts the historical training records, with 5 points for each completed relevant course and an additional 3 points for excellent grades. Finally, it calculates the weighted scores of each item according to the weights: the weight of the work experience score is 30%, the weight of the qualification certificate score is 40%, and the weight of the training record score is 30%, and calculates the final knowledge level score. The score range is 0 - 100 points, which reflects the user's professional knowledge level and serves as an important basis for subsequent knowledge classification.

[0082] S208. Read the job code, job responsibility list, and skill requirement list of the target user from the user profile database, and generate job information parameters including job characteristic identifiers.

[0083] The job code is the unique code that identifies different job positions. For example, "MT001" represents a maintenance technician. The job responsibility list lists the specific work content and requirements of the job. The skill requirement list specifies the professional skills and level requirements for the job. The job characteristic identifier is a standardized description of the key characteristics of the job, including characteristic values in dimensions such as professional field, responsibility scope, and skill requirements. The job information parameter is a structured representation of job characteristics.

[0084] The server extracts job-related information from the user profile database. First, it reads the job code and parses information such as job category and level from the code. Subsequently, it parses the job responsibility list, extracts key responsibility items, and converts them into standardized responsibility codes. Then it analyzes the skill requirement list and maps the skill items to a predefined skill classification system. Finally, it integrates all the information to generate a job characteristic identifier containing a professional field code, a responsibility scope code, and a skill level code. This identifier comprehensively describes the characteristics of the job and provides a benchmark for subsequent knowledge matching.

[0085] S209. Classify the structured knowledge data according to the knowledge level parameter, divide the structured knowledge data into primary, intermediate, and advanced levels, and obtain difficulty-level classified knowledge data.

[0086] Primary, intermediate, and advanced are the three levels of knowledge difficulty, corresponding to basic knowledge, advanced knowledge, and professional knowledge respectively. The structured knowledge data contains the organized professional knowledge content and its attributes. The difficulty-level classified knowledge data is a knowledge system organized by difficulty level. The knowledge level parameter is a quantitative indicator of the user's knowledge level.

[0087] The server classifies the structured knowledge data according to the knowledge level parameter. First, set the difficulty classification criteria: beginner level corresponds to 0 - 60 points, intermediate level corresponds to 61 - 85 points, and advanced level corresponds to 86 - 100 points. Then analyze the characteristics of each knowledge unit in the structured knowledge data, including the complexity of professional terms, the complexity of operation steps, the requirements for prerequisite knowledge, etc. Next, evaluate the difficulty value of each knowledge unit based on these characteristics. Subsequently, allocate the knowledge units to the corresponding difficulty levels according to the difficulty values. Finally, within each difficulty level, organize the knowledge structure according to the association relationship between knowledge units to form a complete difficulty - graded knowledge system.

[0088] S210. Calculate the matching degree between the job feature identifier in the job information parameter and the knowledge label in the difficulty - graded knowledge data, sort the knowledge content according to the matching degree score, and select the knowledge content with a matching degree exceeding the preset matching threshold as the target knowledge data.

[0089] The job feature identifier is a standardized description code for the dimensional characteristics such as the job's professional field, responsibility scope, and skill requirements. The knowledge label is an attribute mark for each knowledge unit in the difficulty - graded knowledge data, containing information such as knowledge type, application scenario, and skill requirements. The matching degree calculation is to calculate the similarity between two sets of identifiers through the vector space model. The matching degree score is a quantitative value of the similarity between identifiers, with a value range of 0 - 1. The preset matching threshold is the standard score for screening knowledge content, usually set to 0.7. The target knowledge data is a set of knowledge content after matching and screening.

[0090] The server performs the matching calculation between the job feature identifier and the knowledge label: First, convert the job feature identifier into a feature vector, with each dimension corresponding to a feature attribute. Subsequently, convert the knowledge label into a vector form as well. Calculate the matching degree score by calculating the cosine similarity of the two vectors. The specific calculation method is: divide the dot product of the two vectors by the product of the magnitudes of their respective vectors. After calculating the matching degree for all knowledge content, sort them from high to low according to the matching degree score. Select the knowledge content with a matching degree score greater than 0.7 to form the target knowledge data set. These knowledge contents have a high correlation with the job requirements and are suitable as training content.

[0091] S211. Select the teaching materials matching the target knowledge data from the knowledge resource library and generate a micro - course resource package based on the teaching materials.

[0092] The knowledge resource library is a data system that stores various teaching materials, including multimedia resources such as text, pictures, videos, and animations. Teaching materials are the carriers of specific teaching content, such as knowledge point explanation documents, operation demonstration videos, exercise questions, etc. The micro-course resource package is a complete learning unit, which includes components such as teaching objectives, learning content, and practice tests. Target knowledge data is the screened knowledge content related to the post.

[0093] The server selects teaching materials from the knowledge resource library and generates micro-courses: First, the target knowledge data is divided into several learning units according to knowledge points. For each learning unit, relevant teaching materials are retrieved from the knowledge resource library, including text explanations, diagram illustrations, video demonstrations, etc. The selected teaching materials need to meet the requirements of complete knowledge point coverage, diverse expression forms, and moderate difficulty levels. The selected materials are organized according to teaching rules to form a standard course structure including five links: "introduction - explanation - example - exercise - summary". The duration of each link is controlled within 5 - 10 minutes to ensure the learning effect. Finally, all resources are packaged to generate a complete micro-course resource package including course information, learning objectives, teaching content, and assessment criteria.

[0094] Next, a more specific process description of the method provided in this implementation will be given. Please refer to Figure 2 , which is another process schematic diagram of the micro-course generation method based on OCR recognition in the embodiments of this application.

[0095] S301. Select teaching materials that match the target knowledge data from the knowledge resource library, and generate a micro-course resource package based on the teaching materials.

[0096] It can be understood that this step is similar to step S105 and will not be elaborated here.

[0097] S302. Based on the fault code information and the operation step information, construct a fault scenario simulation exercise unit.

[0098] The fault code information is standardized data including content such as fault numbers, fault descriptions, and fault types. The operation step information is the specific operation guidance in the maintenance process, including tool usage, parameter settings, operation sequences, etc. The fault scenario simulation exercise unit is a virtual training environment used to simulate maintenance operations in real fault situations, including components such as equipment status display, operation interface, and feedback mechanism.

[0099] The process of building a fault scenario simulation exercise unit for a server first establishes a 3D device model to fully present the external structure and internal components of the device. Then, the fault code information is converted into specific device state manifestations, including abnormal indications, alarm messages, parameter deviations, etc. Next, the operation step information is converted into an interactive operation interface to implement functions such as tool selection, parameter adjustment, component disassembly and assembly. On this basis, a physics engine is added to achieve collision detection and dynamic response between parts. Finally, an acoustic-optic feedback mechanism is set up to give warnings for incorrect operations and confirmations for correct operations. The entire exercise unit is implemented using Web3D technology, supporting multi-angle observation and free operation.

[0100] S303. Set multiple checkpoints in the fault scenario simulation exercise unit, and each checkpoint corresponds to a key operation in the operation step information.

[0101] A checkpoint is a monitoring node set in the operation process to verify the correctness and integrity of the operation. A key operation is a step in the operation steps that plays a decisive role in troubleshooting, such as core parameter adjustment, key component replacement, etc. The purpose of setting checkpoints is to ensure that users complete the repair operation according to the standard process and avoid omitting important steps.

[0102] The specific implementation of setting checkpoints in the server's fault scenario simulation exercise unit first analyzes the operation step information to identify key operation nodes. For each key operation node, trigger conditions and verification rules are set. The trigger conditions include the correctness of tool selection, the rationality of operation timing, the accuracy of parameter setting, etc. The verification rules include operation action determination, parameter range check, status change confirmation, etc. The positions of the checkpoints are marked by highlighting prompts in the 3D scene. Each checkpoint is configured with an independent data acquisition module to record the user's operation data. Logical associations are established between the checkpoints to form a complete operation verification link.

[0103] S304. Record the operation behavior data of the target user at each checkpoint and compare the operation behavior data with the standard operation steps in real time.

[0104] Operation behavior data is all behavior records generated by the user during the operation, including operation selections, operation times, operation sequences, etc. The standard operation steps are pre-defined correct operation processes, including the specific requirements and determination criteria for each step. Real-time comparison is to immediately compare and analyze the user's actual operations with the standard steps.

[0105] The process of the server recording and comparing user operation behaviors first deploys a data collection module at each checkpoint to record all operation information of the user. The collected data includes operation timestamps, tool selection records, operation action sequences, parameter adjustment values, etc. The collected operation behavior data is structured in a predetermined format. At the same time, the standard operation step data is loaded to establish an operation specification comparison table. Through a data comparison algorithm, the differences between the user operation and the standard steps in terms of timing, accuracy, integrity, etc. are analyzed. The comparison results are used to evaluate the correctness of the operation and serve as the basis for the next feedback and guidance.

[0106] S305. When a wrong operation of the target user is detected, extract the professional term information and operation step information related to the standard operation step from the target knowledge data to generate prompt content, and send it to the target client.

[0107] A wrong operation refers to a behavior that deviates from the standard operation steps, including situations such as incorrect operation sequence, incorrect tool selection, incorrect parameter setting, etc. Professional term information is the explanation of professional terms related to the operation, including term definitions, usage scenarios, precautions, etc. Operation step information includes specific operation guides, such as tool usage methods, parameter adjustment ranges, operation precautions, etc. Prompt content is the guidance information generated for the wrong operation, including error reminders, correct demonstrations, and improvement suggestions. The target client is the terminal device used by the user for operation practice.

[0108] The processing process after the server detects a wrong operation first identifies the type of error, compares the user's actual operation with the standard steps to determine the link and specific manifestation where the error occurs. Retrieve the knowledge content related to the current operation step from the target knowledge data, and extract the corresponding professional term explanations and detailed operation instructions. Generate targeted prompt content based on the extracted information, and the content includes three parts: error description, professional knowledge supplement, and correct operation guidance. Push the generated prompt content to the target client in real time through WebSocket and display it in a pop-up window or highlighted form at an appropriate position on the interface. The display duration of the prompt content is set to 15 seconds and will automatically close after the user confirms.

[0109] S306. Obtain the score of the target user in the fault scenario simulation practice unit.

[0110] The score situation is a quantitative evaluation of the user's performance during the practice, including indicators such as operation accuracy rate, completion time, and number of errors. The scoring in the fault scenario simulation practice unit is based on preset scoring rules to score and count each operation of the user. The score data is used to evaluate the user's learning effect and operation proficiency.

[0111] The specific implementation process for the server to obtain the scoring situation is as follows: First, it counts the operation data of the user at each checkpoint, including operation correctness, operation time, number of retries, etc. Different score weights are set for each checkpoint, and the weights for important operation steps are higher. The basic score is calculated based on the accuracy of the operation: a correct operation gets full marks, a wrong operation that is corrected after a prompt gets partial marks, and multiple wrong operations get no marks. Considering the operation time factor, points are added for completion within the standard time, and points are deducted for overtime. The number of wrong operations is counted, and corresponding points are deducted each time. Finally, the scores of each item are summarized to generate a detailed scoring report including the total score and the scores of each sub-item.

[0112] S307. Adjust the difficulty level and knowledge point distribution of the micro-course resource package according to the scoring situation.

[0113] The difficulty level of the micro-course resource package is divided into three levels: elementary, intermediate, and advanced, and each level corresponds to knowledge content with different depths. The knowledge point distribution refers to the arrangement and proportion of knowledge content in the course. The difficulty adjustment means adjusting the depth and breadth of the teaching content according to the user's performance. The knowledge point distribution adjustment means optimizing the arrangement order and explanation details of the knowledge content.

[0114] The specific implementation for the server to adjust the micro-course according to the scoring situation is as follows: First, it analyzes the score distribution of the user on different knowledge points and identifies the knowledge points that are poorly mastered. The adjustment direction is determined according to the score level: when the total score is lower than 60, the difficulty level is reduced and the basic knowledge explanation is supplemented; when the total score is between 60 and 85, the current difficulty is maintained and the weak knowledge points are strengthened; when the total score is higher than 85, the difficulty level is increased and in-depth content is added. When adjusting the knowledge point distribution, the explanation space and the number of practice times for the knowledge points with lower scores are increased, and the knowledge points that are well mastered are appropriately simplified. The order of the course content is reorganized to ensure reasonable connection of knowledge points and smooth progression of difficulty. A new micro-course resource package is generated, and the course outline, learning objectives, and assessment criteria are updated.

[0115] S308. Statistically analyze the learning behavior data of multiple target users using the same micro-course resource package, and extract the knowledge point access frequency, average learning duration, and number of repeated learning times from the learning behavior data.

[0116] Learning behavior data is a data set that records various operations of users during the learning process, including access records, stay time, number of repetitions, and other information. The knowledge point access frequency refers to the statistical count of the number of times a specific knowledge content is viewed by the user. The average learning duration is the average learning time of the user on each knowledge point. The number of repeated learning times is the statistical count of the user's repeated learning behavior for the same knowledge point. Statistical analysis is a process of performing numerical calculations and extracting rules from these data.

[0117] The statistical analysis process of the server for learning behavior data first extracts all the learning records of target users from the database. It counts the access volume of each knowledge point, records the access timestamp and access duration. It calculates the daily average access volume and total access volume of each knowledge point to obtain the knowledge point access frequency data. It analyzes the residence time of users on each knowledge point, and calculates the average learning duration after removing outliers. It identifies the repeated access behavior of users to knowledge points, counts the number of times a single user learns the same knowledge point, and calculates the repeated learning times. It generates a statistical report including access frequency distribution, duration distribution, and repetition rate distribution.

[0118] S309. Perform importance annotation on the technical term information and operation step information in the structured knowledge data according to the access frequency of this knowledge point, this average learning duration, and this repeated learning times to obtain annotation information.

[0119] Importance annotation is a quantitative identification of the value degree of knowledge content. Technical term information is technical terms and their explanatory notes. Operation step information is the execution guidance of specific operations. Annotation information is a numerical representation of the importance degree of knowledge content, which is used to guide the subsequent optimization of content presentation. Structured knowledge data is a set of knowledge content that has been systematically sorted out.

[0120] The specific implementation of the server for performing importance annotation based on statistical data first establishes an annotation scoring system, and sets the weight ratios of three dimensions: access frequency, learning duration, and repetition times. It normalizes the access frequency data of each knowledge point and converts it into a standard score of 0-100. It calculates the ratio of the average learning duration to the preset standard duration to generate a quantitative score for the duration dimension. It converts the repeated learning times into a repetition rate index and calculates the repeated learning tendency score. It comprehensively calculates the overall importance score of the knowledge point according to the scores of the three dimensions and the weight ratios. It performs hierarchical annotation on the knowledge content according to the scores, and divides it into three levels: core content, important content, and general content. It generates an annotation data set including knowledge point numbers, importance levels, and scores of each dimension.

[0121] S310. Adjust the presentation order and display duration of teaching materials in the micro-course resource package based on this annotation information.

[0122] Teaching materials are the specific carriers of knowledge content, including forms such as text, pictures, and videos. Presentation order is the arrangement order of teaching materials in the course. Display duration is the playing or display time of each teaching material. Annotation information is a quantitative indicator of the importance of knowledge content. Micro-course resource package is a complete learning course unit.

[0123] The specific implementation of the server to adjust teaching materials based on annotation information first reads the importance level and score data in the annotation information. Sort the teaching materials according to the importance level, place the core content at the beginning of the unit, the important content follows immediately, and the general content is placed in the subsequent positions. Adjust the display duration of each teaching material, extend the display time for the core content, and appropriately compress the duration for the general content. Rearrange the association order between knowledge points to ensure reasonable connection of the content before and after. Optimize the presentation form of the teaching materials, add graphical explanations and detailed interpretations to the important content. Generate a new material arrangement plan and update the content order and time arrangement in the micro-course resource package.

[0124] S311. Collect the problem feedback submitted by the target user during the learning process, and establish an association between the problem points that appear more than the preset frequency and the target knowledge content in the target knowledge data.

[0125] Problem feedback is the doubts, suggestions, and difficulty descriptions put forward by users during the learning process, recorded in text form. The preset frequency is the quantitative standard for judging the importance of problems, usually set to 10 times per month. Problem points are the specific knowledge or operation difficulties extracted from the problem feedback. Target knowledge content is the explanatory materials in the knowledge base related to the problem points. The association relationship is the mapping connection established between the problem points and the solutions.

[0126] The specific implementation of the server to collect and process problem feedback first collects the problem content submitted by users through the Q&A interface. Perform natural language processing on the collected problem text to extract keywords and core problem descriptions. Classify and aggregate according to the problem theme, and count the submission frequency of each type of problem. Screen out the problems with a monthly submission frequency exceeding 10 times and determine them as key problems. Retrieve the knowledge content related to the key problems from the target knowledge data, including concept explanations, operation instructions, precautions, etc. Establish a mapping table from problem points to knowledge content, and record information such as problem ID, knowledge point ID, and association degree. Store these association relationships in the database as the basis for subsequent content optimization.

[0127] S312. Update the content structure of the micro-course resource package according to the target knowledge content.

[0128] The content structure of the micro-course resource package is the organizational framework of teaching content, including elements such as knowledge unit division, content order arrangement, and key point and difficulty marking. Target knowledge content is the knowledge point materials that need to be emphasized. Content structure update refers to optimizing the organization method of course content based on learning feedback.

[0129] The specific implementation of the server to update the micro-course content structure first loads the original course structure data and identifies the positions of the knowledge points that need to be strengthened. Add detailed concept explanations and example illustrations at the relevant knowledge points to expand the depth and breadth of the learning materials. Adjust the presentation method of the knowledge points, and add chart explanations and interactive examples to the key content. Rearrange the order of the knowledge units, and explain the knowledge points corresponding to the high-frequency questions in advance. Insert targeted exercises and self-test questions into the course to strengthen the consolidation of key knowledge. Update the course navigation and knowledge map to highlight the position markers of the key content. Generate a new version of the course structure file, and update the course outline and content index in the resource package. Record the updated version number and modified content for subsequent tracking and evaluation.

[0130] The server in the embodiment of the present invention application will be described below from the perspective of hardware processing. Please refer to Figure 4 , which is a schematic structural diagram of an entity device of the server in the embodiment of the present application.

[0131] It should be noted that Figure 4 The structure of the server shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.

[0132] As Figure 4 shown, the server includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 402 or the program loaded from the storage part 408 into the random access memory (RAM) 403, such as executing the method described in the above embodiments. In the RAM 403, various programs and data required for system operation are also stored. The CPU 401, ROM 402, and RAM 403 are connected to each other through a bus 404. The input / output (I / O) interface 405 is also connected to the bus 404.

[0133] The following components are connected to the I / O interface 405: an input section 406 including an audio input device, a button switch, etc.; an output section 407 including a liquid crystal display (LCD), an audio output device, an indicator light, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the I / O interface 405 as required. A removable medium 411 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is installed on the drive 410 as required so that a computer program read therefrom can be installed into the storage section 408 as required.

[0134] Specifically, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 409, and / or installed from the removable medium 411. When the computer program is executed by a central processing unit (CPU) 401, various functions defined in the present invention are executed.

[0135] It should be noted that specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, device, or apparatus.

[0136] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the block may occur in a different order from that marked in the accompanying drawings.

[0137] Specifically, the server of this embodiment includes a processor and a memory, and a computer program is stored on the memory. When the computer program is executed by the processor, it implements the method for generating micro-courses based on OCR recognition provided in the above-mentioned embodiment.

[0138] On the other hand, the present invention also provides a computer-readable storage medium, which may be included in the server described in the above-mentioned embodiment; or it may exist separately without being assembled into the server. The above storage medium carries one or more computer programs. When the above one or more computer programs are executed by a processor of the server, the server is enabled to implement the method for generating micro-courses based on OCR recognition provided in the above-mentioned embodiment.

[0139] As mentioned above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the various embodiments of the present application.

[0140] As used in the above embodiments, according to the context, the term "when..." can be interpreted to mean "if...", "after...", "in response to determining...", or "in response to detecting...". Similarly, according to the context, the phrase "when determining..." or "if detecting (the stated condition or event)" can be interpreted to mean "if determining...", "in response to determining...", "when detecting (the stated condition or event)", or "in response to detecting (the stated condition or event)".

[0141] Those of ordinary skill in the art can understand all or part of the processes in the methods of the above embodiments. These processes can be completed by relevant hardware instructed by a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. The foregoing storage medium includes various media that can store program codes, such as ROM or random access memory RAM, magnetic disks, or optical discs.

Claims

1. A micro-course generation method based on OCR recognition, characterized in that: Applied to a server, the method comprises: Collecting maintenance record document image data, performing optical character recognition processing on the maintenance record document image data to obtain digital text data; Inputting the digital text data into a preset bidirectional encoder representation model to generate a text vector containing contextual semantic information; Extracting professional terminology information, fault code information and operation step information from the text vector; Establishing a first mapping relationship between the fault code information and the professional terminology information, and a second mapping relationship between the fault code information and the operation step information, and determining fault knowledge association data based on the first mapping relationship and the second mapping relationship; Classify and organize the fault knowledge-related data, and establish a hierarchical structure including fault types, fault characteristics and solutions according to the fault frequency, maintenance difficulty and equipment importance; Connecting the hierarchical structures in series according to the order of fault processing in the solution to form structured knowledge data; Acquire knowledge level parameters and job information parameters of the target user, and perform difficulty grading on the structured knowledge data according to the knowledge level parameters to obtain difficulty graded knowledge data; Screening the knowledge range in the difficulty graded knowledge data according to the job information parameters to obtain target knowledge data; Select teaching materials that match the target knowledge data from the knowledge resource library, and generate a micro-course resource package based on the teaching materials.

2. The method according to claim 1, characterized in that The steps of obtaining the knowledge level parameter and the job information parameter of the target user, grading the difficulty of the structured knowledge data according to the knowledge level parameter to obtain the difficulty graded knowledge data, and screening the knowledge range in the difficulty graded knowledge data according to the job information parameter to obtain the target knowledge data specifically include: Read the target user's working years data, professional qualification certificate data and historical training record data from the user archive database, calculate the knowledge level score according to the pre-designed scoring rules, and use the knowledge level score as the knowledge level parameter; Reading the target user's job code, job responsibilities list and skill requirement list from the user profile database, and generating job information parameters including job feature identifiers; According to the knowledge level parameter, the structured knowledge data is graded in difficulty, and the structured knowledge data is divided into elementary, intermediate and advanced levels to obtain knowledge data with graded difficulty; The matching degree between the position feature identifier in the position information parameter and the knowledge label in the difficulty graded knowledge data is calculated, the knowledge contents are sorted according to the matching degree score, and the knowledge contents whose matching degree exceeds a preset matching threshold are selected as the target knowledge data.

3. The method according to claim 1, characterized in that After the step of selecting teaching materials matching the target knowledge data from the knowledge resource library and generating a micro-course resource package according to the teaching materials, the method further includes: Based on the fault code information and the operation step information, construct a fault scenario simulation exercise unit; Setting a plurality of checkpoints in the fault scenario simulation exercise unit, each of the checkpoints corresponding to a key operation in the operation step information; Recording the target user's operation behavior data at each of the checkpoints, and comparing the operation behavior data with the standard operation steps in real time; When an erroneous operation of the target user is detected, the professional terminology information and operation step information related to the standard operation steps are extracted from the target knowledge data to generate prompt content, and sent to the target client.

4. The method according to claim 3, characterized in that After the step of extracting professional terminology information and operation step information related to the standard operation steps from the target knowledge data to generate prompt content when an erroneous operation of the target user is detected, the method further includes: Obtaining the score of the target user in the fault scenario simulation practice unit; The difficulty level and knowledge point distribution of the micro-course resource package are adjusted according to the score.

5. The method according to claim 1, characterized in that After the step of selecting teaching materials matching the target knowledge data from the knowledge resource library and generating a micro-course resource package according to the teaching materials, the method further includes: Performing statistical analysis on the learning behavior data of multiple target users using the same micro-course resource package, and extracting the knowledge point access frequency, average learning time and number of repeated learning in the learning behavior data; According to the knowledge point access frequency, the average learning time and the number of repeated learning times, importance marking is performed on the professional terminology information and the operation step information in the structured knowledge data to obtain marking information; The presentation order and display duration of the teaching materials in the micro-course resource package are adjusted based on the annotation information.

6. The method according to claim 1, characterized in that After the step of selecting teaching materials matching the target knowledge data from the knowledge resource library and generating a micro-course resource package according to the teaching materials, the method further includes: Collecting the problem feedback submitted by the target user during the learning process, and associating the problem points that appear more than a preset frequency with the target knowledge content in the target knowledge data; The content structure of the micro-course resource package is updated according to the target knowledge content.

7. A server, characterized in that: The server includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the server to execute the method described in any one of claims 1-6.

8. A computer-readable storage medium comprising instructions, characterized in that: When the instructions are executed on a server, the server is caused to execute the method according to any one of claims 1 to 6.

9. A computer program product, characterized in that When the computer program product is run on a server, the server is caused to execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Equipment maintenance knowledge determination method and device, equipment and storage medium

    CN117764551A

  • Course editing method of rail transit weak current equipment training system

    CN118734798A