Practical operation detection result information generation system
By using a head-mounted display device and processing system, and leveraging SLAM technology and speech recognition, the problem of low detection accuracy in existing technologies has been solved, enabling efficient and accurate generation of practical detection results.
Patent Information
- Application Number
- CN202511612639.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-05
- Publication Date
- 2026-02-10
AI Technical Summary
Existing practical testing technologies cannot effectively capture and detect key skill indicators such as the standardization of gestures, force control, and spatial trajectory accuracy involved in users' actual operations, resulting in low accuracy of test results. It is necessary to introduce additional video recording, which increases time consumption and computing resources.
By employing a head-mounted display device and a processing system, a real-world detection scenario map is generated using SLAM technology. Combined with voice and operational action information, action and voice recognition results are generated to improve detection accuracy.
It improved the accuracy of practical test results, reduced generation time and computational resource consumption, and enhanced the test effect.
Smart Images

Figure CN121504685A_ABST
Abstract
Description
Technical Field
[0001] The embodiments disclosed herein relate to the field of computer technology, and more specifically to a system for generating practical test result information. Background Technology
[0002] Practical skills assessment is a core means of evaluating the skill level, operational standardization, and problem-solving abilities of trainees or employees, and it is widely used in many fields such as vocational education, military training, medical surgery, industrial maintenance, and emergency drills. Currently, the common method for implementing practical skills assessment is to deploy personal computers in a fixed examination room, where users interact with virtual teaching aid models displayed on the screen using a mouse, keyboard, or touchscreen to simulate operations.
[0003] However, in practice, it has been found that when using the above method to implement practical testing, the following technical problems often arise: The interaction logic relying on common input devices such as mice and keyboards differs from professional skills such as hand operations and tool use in real-world environments. It cannot capture and detect key skill indicators such as the standardization of gestures, force control, and spatial trajectory accuracy involved in real operations, resulting in a disconnect between the detected content and the practical scenario, and poor effectiveness in detecting user operation skills. The system can only record simple interaction events (such as click coordinates and operation sequences), resulting in low accuracy of the practical operation detection results. It is necessary to introduce additional video recording and review of the videos, which increases the time and computational resources required to generate practical operation detection results.
[0004] The information disclosed in this background section is only intended to enhance the understanding of the background of the inventive concept, and therefore may contain information that does not form prior art known to those skilled in the art. Summary of the Invention
[0005] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.
[0006] Some embodiments of this disclosure propose a practical test result information generation system to solve one or more of the technical problems mentioned in the background section above.
[0007] In a first aspect, some embodiments of this disclosure provide a practical operation test result information generation system, including an authoring end, a head-mounted display device, and a processing end; the authoring end is used to obtain map information of the current practical operation test scenario from the processing end; based on the map information, it edits the practical operation test standard information of the current practical operation test scenario to obtain test standard editing information; the head-mounted display device is used to generate an action judgment result information set according to the test standard editing information and the collected operation action information set; the processing end is used to generate a speech recognition result information set according to the collected speech information set and the test standard editing information; and generate practical operation test result information according to the action judgment result information set and the speech recognition result information set.
[0008] Optionally, the method further includes: a map acquisition terminal, used to generate map information corresponding to each of the above-mentioned practical testing teaching aids based on the scanning result information of each practical testing teaching aid; and uploading the map information to the above-mentioned processing terminal.
[0009] Optionally, the method further includes: a head-mounted display device, configured to perform the following determination steps based on each operation action information in the above-mentioned operation action information set: in response to determining that the above-mentioned operation action information meets a preset determination condition, correctly determining the action detection information as action determination result information; in response to determining that the above-mentioned operation action information does not meet the above-mentioned preset determination condition, incorrectly determining the action detection information as action determination result information; and determining the obtained action determination result information as a set of action determination result information.
[0010] Optionally, the method further includes: a head-mounted display device, used to determine process log information based on a set of text information corresponding to the aforementioned voice information set, a coordinate trajectory corresponding to the aforementioned hand position information, and a timestamp; and to upload the aforementioned process log information to the aforementioned processing terminal.
[0011] Optionally, the method further includes: a processing end, configured to perform the following recognition steps for each piece of speech information in the aforementioned speech information set: matching the speech information according to the aforementioned detection standard editing information to obtain matching degree information; correctly identifying the speech information as speech recognition result information in response to determining that the aforementioned matching degree information meets a preset threshold condition; incorrectly identifying the speech information as speech recognition result information in response to determining that the aforementioned matching degree information does not meet the aforementioned preset threshold condition; and determining the obtained speech recognition result information set.
[0012] Secondly, some embodiments of this disclosure provide a method for generating practical test result information, applied to a head-mounted display device included in the practical test result information generation system described in the first aspect above, comprising: collecting a set of voice information and a set of operation action information; editing information according to the above-mentioned test standards and the above-mentioned set of operation action information to generate a set of action judgment result information; and sending the above-mentioned set of voice information and the above-mentioned set of action judgment result information to a processing terminal included in the above-mentioned practical test result information generation system.
[0013] Thirdly, some embodiments of this disclosure provide a head-mounted display device, including: one or more processors; a storage device for storing one or more programs; an optomechanical system and optical elements for imaging in front of a user; and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation of the eighth aspect above.
[0014] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method as described in any implementation of the eighth aspect.
[0015] The above-described embodiments of this disclosure have the following beneficial effects: The practical operation test result information generation system of some embodiments of this disclosure can improve the effectiveness of user operation skill testing, increase the accuracy of practical operation test result information, and reduce the time and computational resources required to generate practical operation test result information. Specifically, the reasons for the low validity of the assessment results and the increased time and computational resources required to generate practical operation test result information are as follows: the interaction logic relying on general input devices such as mice and keyboards differs from professional skills such as hand operations and tool usage in real-world environments. It cannot capture and detect key skill indicators such as gesture standardization, force control, and spatial trajectory accuracy involved in real-world operations, resulting in a disconnect between the detected content and the practical scenario, leading to poor effectiveness in testing user operation skills. Furthermore, the system can only record simple interaction events (such as click coordinates and operation sequences), resulting in low accuracy of the obtained practical operation test result information. This necessitates additional video recording and review of the videos, further increasing the time and computational resources required to generate practical operation test result information. Based on this, the practical operation test result information generation system of some embodiments of this disclosure includes a creation terminal, a head-mounted display device, and a processing terminal. The creation terminal is used to obtain map information of the current practical operation test scenario from the processing terminal. This allows for the acquisition of various teaching aids and their environments corresponding to the current practical operation test scenario. The creation terminal is used to edit the practical operation test standard information of the current practical operation test scenario based on the map information, obtaining test standard editing information. This allows for the acquisition of assessment standards corresponding to each test point in the current practical operation test scenario. The head-mounted display device is used to generate an action judgment result information set based on the above-mentioned test standard editing information and the collected set of operation action information. This allows for the acquisition of a judgment result for each operation action information in the operation action information set. The processing terminal is used to generate a speech recognition result information set based on the collected set of speech information and the above-mentioned test standard editing information. This allows for the acquisition of a recognition result for each speech information in the speech information set. The processing terminal is used to generate practical operation test result information based on the above-mentioned action judgment result information set and the above-mentioned speech recognition result information set. This allows for the acquisition of the final score after the assessment. Because it doesn't rely on common input devices like mice and keyboards for interaction when generating practical operation test results, but instead simulates practical operation tests in a real-world environment within augmented reality, it can effectively capture and test key skill indicators such as gesture standardization, force control, and spatial trajectory accuracy involved in real-world operations, thus improving the effectiveness of the test results.Furthermore, because the system doesn't just record simple interaction events (such as click coordinates or operation sequences), but also synchronously collects voice information and spatial positioning data provided by SLAM, it can improve the accuracy of practical operation test results. This eliminates the need for additional video recording and review, reducing the time and computational resources required to generate practical operation test results. Consequently, it can improve the validity of the assessment results, increase the accuracy of practical operation test results, and reduce the time and computational resources required to generate practical operation test results. Attached Figure Description
[0016] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.
[0017] Figure 1 This is a schematic diagram of the system for generating information based on the practical test results disclosed herein; Figure 2 This is an exemplary system architecture diagram of a practical test result information generation system applying some embodiments of the present disclosure; Figure 3 These are timing diagrams of some embodiments of the system for generating information from practical test results based on this disclosure; Figure 4 These are flowcharts of some embodiments of the method for generating practical test result information for head-mounted display devices according to the present disclosure; Figure 5 This is a schematic diagram of the hardware structure of a head-mounted display device suitable for implementing some embodiments of the present disclosure. Detailed Implementation
[0018] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0019] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.
[0020] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0021] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0022] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0023] Figure 1 The diagram shows structural schematics of some embodiments of the practical test result information generation system that can be applied to this disclosure.
[0024] like Figure 1 As shown, the practical testing result information generation system provided in this disclosure may include: an authoring terminal 101, a processing terminal 102, and a head-mounted display device 21. The authoring terminal 101 and the processing terminal 102 are communicatively connected. The head-mounted display device 21 and the processing terminal 102 are also communicatively connected. It should be noted that the communication connection may include, but is not limited to, 3G / 4G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra-wideband) connection, and other currently known or future-developed communication methods. The authoring terminal 101 may represent a server used to obtain map information of the current practical testing scenario from the processing terminal; based on the map information, it updates the practical testing standard information of the current practical testing scenario to obtain testing standard editing information. The processing terminal 102 may represent a cloud connection. The processing terminal 102 can be used to generate a speech recognition result information set based on the collected speech information set and the testing standard editing information; and to generate practical testing result information based on the action judgment result information set and the speech recognition result information set. The aforementioned practical testing result information generation system also includes terminal device 22 and a map acquisition terminal. The map acquisition terminal can represent a device used to scan and create a map of the practical testing scenario to obtain map information. The map acquisition terminal can represent a mobile device with SLAM (Simultaneous Localization and Mapping) functionality. The specific type of the mobile device is not limited; for example, the mobile device can be a mobile phone.
[0025] like Figure 2 As shown above, Figure 2 An exemplary system architecture 200 can characterize the aforementioned head-mounted display device 21 and the aforementioned terminal device 22. Figure 2 It may include the head-mounted display device 21 and the terminal device 22 described above.
[0026] The head-mounted display device 21 may include one or two optomechanical and optical elements 211. These optomechanical and optical elements are used to display a virtual interface. Furthermore, the head-mounted display device 21 also includes a frame 212. In some embodiments, the sensors, processing unit, memory, and battery of the head-mounted display device 21 may be housed inside the frame 212. In some alternative implementations, one or more components of the sensors, processing unit, memory, and battery may be integrated into another separate accessory (not shown), connected to the frame 212 via a data cable 23. In some alternative implementations, the head-mounted display device 21 may only have display functionality and some sensors, while data processing, data storage, and power supply capabilities are provided through a terminal device 22.
[0027] The terminal device 22 includes a touch component 221, which can be a touch screen or a touchpad. The touch screen can be used to display content and control window dragging, and the touchpad can be used to control window dragging. In some embodiments, the head-mounted display device 21 and the terminal device 22 can be connected via a data cable 23.
[0028] It should be understood that Figure 2 The number of head-mounted display devices and terminal devices shown is merely illustrative. Any suitable number of head-mounted display devices and terminal devices can be included depending on the implementation requirements.
[0029] The following is for reference. Figure 3 The diagram shows a timing diagram of a practical test result information generation system of this disclosure.
[0030] like Figure 3 As shown, a practical testing result information generation system includes: a creation terminal, a processing terminal, and a head-mounted display device. The interaction steps between the creation terminal, the processing terminal, and the head-mounted display device may include the following steps: Step 301: The authoring device is used to obtain map information of the current practical testing scenario.
[0031] In some embodiments, the authoring client can obtain map information of the current practical testing scenario. This current practical testing scenario can represent the three-dimensional space in which the user performs the practical testing. For example, the current practical testing scenario can represent a laboratory containing experimental equipment, such as an operating table. The current practical testing scenario can correspond to a set of knowledge points. This set of knowledge points can include multiple knowledge points to be examined. For example, if the current practical testing scenario is a CNC machine tool operation workshop, the knowledge points to be examined could be "machine tool idle operation," "manual control of trial cutting end face," or "timely pause after external diameter trial cutting to avoid over-cutting." The map information can represent a SLAM map of the current practical testing scenario. In practice, the authoring client can obtain the map information of the current practical testing scenario from the processing client.
[0032] Optionally, prior to step 301, the above-mentioned practical test result information generation system further includes a map acquisition terminal, which is configured to perform the following steps: The first step is to generate map information for each of the practical testing aids in the current practical testing scenario, based on the scanning results. Each practical testing aid represents the tool or equipment used during the actual operation assessment. The scanning results represent the 3D point cloud obtained after scanning each practical testing aid. In practice, firstly, the map acquisition end can scan each practical testing aid using a Simultaneous Localization and Mapping (SLAM) acquisition tool to obtain scanning results. Then, using the SLAM acquisition tool, based on the scanning results, a map scan is performed on the environment where each practical testing aid is located to create the corresponding map information.
[0033] The second step is to upload the map information to the processing terminal.
[0034] Step 302: The authoring end is used to edit the practical testing standard information of the current practical testing scenario based on map information to obtain the testing standard editing information.
[0035] In some embodiments, the authoring entity can edit the practical testing standard information of the current practical testing scenario based on the aforementioned map information to obtain testing standard editing information. The practical testing standard information can represent the initial assessment standard set corresponding to the aforementioned knowledge point set. The initial assessment standards in the initial assessment standard set can represent the originally set assessment standards. The originally set assessment standards can be empty. The assessment standards can represent voice test point data or action test point data during practical testing. For example, the assessment standard can be "measure the user's neck length and adjust the neck brace." The voice test point data can represent the standard answer text for the corresponding test point. The action test point data can represent the standard gesture and preset position (three-dimensional spatial coordinates) for the corresponding test point. The preset position can represent the required hand position. The testing standard editing information can represent the modified assessment standard set corresponding to the aforementioned initial assessment standard set. The modified assessment standards in the modified assessment standard set can correspond one-to-one with the knowledge points in the aforementioned knowledge point set. In practice, the aforementioned creation platform allows users to input parameters through a graphical interface to edit the initial assessment standard set represented by the practical testing standard information. The resulting edited initial assessment standard set serves as the modified assessment standard set, and the modified assessment standard set is used as the testing standard editing information. These parameters can represent the pre-set answer text, pre-set gestures, and preset positions for each knowledge point in the aforementioned knowledge point set.
[0036] Step 303: The head-mounted display device is used to generate a set of action judgment results information based on the information edited according to the detection standards and the set of collected operation action information.
[0037] In some embodiments, the head-mounted display device can generate a set of action judgment results information based on the aforementioned standard editing information and the collected set of operation action information. The operation action information in the aforementioned set of operation action information can represent the user's gesture recognition information and hand position information during practical testing (e.g., practical assessment). The gesture recognition information can represent the user's gestures during practical testing (practical assessment). The hand position information can represent the position of the user's hands during practical testing (practical assessment). The action judgment result information in the aforementioned set of action judgment results information can represent the detection result of the operation action information. The detection result can represent whether the operation action information is correct or incorrect. The action judgment result information in the aforementioned set of action judgment results information can have a one-to-one correspondence with the operation action information in the aforementioned set of operation action information. Each operation action information in the aforementioned set of operation action information corresponds to one knowledge point in the aforementioned set of knowledge points.
[0038] Optionally, prior to step 303 above, the head-mounted display device can perform real-time recognition of the target user's hands using a SLAM camera to obtain gesture recognition information. Then, the SLAM camera is used to locate the hand using this gesture recognition information to obtain hand position information. Finally, the obtained hand position information and gesture recognition information are used to determine the operation action information. The target user can be a user currently wearing the head-mounted display device for practical testing. For example, the target user could be a maintenance worker.
[0039] In some optional implementations of certain embodiments, the head-mounted display device may be further configured to perform the following steps to generate an action determination result information set based on the above-described standard editing information and the collected set of operation action information: The first step is to perform the following determination steps based on each operation action information in the above set of operation action information: The first sub-step involves correctly determining the action detection information as the action judgment result information in response to determining that the aforementioned action information meets preset judgment conditions. Here, "correct" action detection information indicates that the detection result of the aforementioned action information is correct. The preset judgment conditions can be that the gesture recognition information included in the aforementioned action information meets preset standard conditions, and the hand position information included in the aforementioned action information meets preset error threshold conditions. The preset standard conditions can be that the gesture recognition information is consistent with the standard gesture included in the corresponding changed assessment standard. The changed assessment standard corresponding to the aforementioned gesture recognition information can be a changed assessment standard where the corresponding knowledge point is the same as the knowledge point corresponding to the aforementioned gesture recognition information. The preset error threshold condition can be that the distance between the aforementioned hand position information and the preset position included in the corresponding changed assessment standard is within a preset range. The changed assessment standard corresponding to the aforementioned hand position information can be a changed assessment standard where the corresponding knowledge point is the same as the knowledge point corresponding to the aforementioned hand position information. Here, the specific range of the aforementioned preset range is not limited. For example, the preset range can be -5cm to +5cm.
[0040] The second sub-step involves determining that the aforementioned operation action information does not meet the aforementioned preset judgment conditions, and classifying the action detection information as an error as the action judgment result information. Here, the aforementioned action detection information error indicates that the detection result of the aforementioned operation action information is incorrect.
[0041] The second step is to determine the obtained action judgment result information into a set of action judgment result information.
[0042] Step 304: The processing end is used to generate a speech recognition result information set based on the collected speech information set and the detection standard editing information.
[0043] In some embodiments, the processing unit can generate a speech recognition result information set based on the collected speech information set and the aforementioned detection criteria editing information. The speech information in the aforementioned speech information set can represent the statements given by the user wearing the aforementioned head-mounted display device during practical testing. For example, the speech information could be "The machine tool worktable is free of debris, the guide rail protective cover is intact, and all operating handles are in neutral." The speech recognition result information in the aforementioned speech recognition result information set can represent the recognition result corresponding to the speech information in the aforementioned speech information set. The recognition result can indicate whether the aforementioned speech information is correct or incorrect.
[0044] In some optional implementations of certain embodiments, the processing end may be further configured to perform the following steps to generate a speech recognition result information set based on the collected speech information set and the above-mentioned detection criteria: The first step is to perform the following recognition steps for each piece of speech information in the above speech information set: The first sub-step involves editing the information according to the aforementioned detection criteria and performing matching processing on the aforementioned speech information to obtain matching degree information. This matching degree information characterizes the degree of matching between the text corresponding to the aforementioned speech information and the corresponding changed assessment criteria. The changed assessment criteria corresponding to the aforementioned text characterize the changed assessment criteria where the corresponding knowledge point is the same as the knowledge point corresponding to the aforementioned text. In practice, firstly, the processing end can convert the aforementioned speech information into text using ASR technology. Then, using a word vector model, the matching degree between the aforementioned text and the corresponding changed assessment criteria is determined. Finally, the aforementioned matching degree is used as the matching degree information. The aforementioned ASR technology can be Alibaba Cloud Speech Recognition. The aforementioned word vector model can be Word2Vec.
[0045] The second sub-step involves correctly identifying the speech information as a speech recognition result in response to determining that the matching degree information meets a preset threshold condition. The preset threshold condition can be that the matching degree information is greater than or equal to a preset value. The specific value of the preset value is not limited here; for example, the preset value could be 90%. The statement "the speech information is correct" indicates that the speech information recognition result is correct.
[0046] The third sub-step involves determining that the matching degree information does not meet the preset threshold condition, and classifying the voice information as an error in the voice recognition result. Here, the aforementioned voice information error can indicate that the recognition result of the voice information is incorrect.
[0047] The second step is to determine the obtained speech recognition results into a speech recognition result set.
[0048] Optionally, before step 304, the head-mounted display device can send the voice information set and the action determination result information set to the processing terminal. In practice, the head-mounted display device can send the voice information set and the action determination result information set to the processing terminal via the processing terminal's IP address.
[0049] Step 305: The processing end generates practical operation detection result information based on the action judgment result information set and the speech recognition result information set.
[0050] In some embodiments, the processing end can generate practical operation detection result information based on the aforementioned action determination result information set and the aforementioned speech recognition result information set. The practical operation detection result information can represent the final detection result corresponding to the practical operation detection. The detection result can represent the user's final score. For example, the final detection result can be 80 points. In practice, firstly, the processing end can determine the total number of action determination result information in the aforementioned action determination result information set representing correct detection results and the total number of speech recognition result information in the aforementioned speech recognition result information set as a first total. Then, the total number of action determination result information in the aforementioned action determination result information set and the total number of speech recognition result information in the aforementioned speech recognition result information set is determined as a second total. Next, the ratio between the first total and the second total is determined as a target value. Finally, the product of the target value and 100 is determined as the practical operation detection result information.
[0051] In addressing the technical problems mentioned above, and considering the application scenario of practical operation on CNC lathes (milling machines), the following technical problem often arises: using a fixed error threshold leads to low evaluation accuracy and insufficient discrimination, increasing the amount of invalid assessment data to be processed. This forces the system to use more questions and more time to evaluate candidates' abilities to obtain the final practical test results, thus reducing the processing efficiency and resource utilization of practical test results. Given the following requirements for this application scenario—high precision and complexity of operations—we have decided to adopt the following solution: Optionally, after step 305, the aforementioned processing terminal can be further configured to perform the following steps: The first step is to acquire a first set of operation action information collected from the head-mounted display device. The first operation action information in this set represents the user's gestures and hand positions during the practical operation test. The specific number of first operation action information items in this set is not limited. For example, the number of first operation action information items in the set can be 5. It should be noted that the method of acquiring the first operation action information in the head-mounted display device's first set of operation action information is the same as the method of acquiring operation action information in the head-mounted display device's operation action information set, and will not be repeated here; please refer to step 303.
[0052] The second step involves determining the aforementioned preset error threshold condition as the target error threshold condition, the preset number of rounds as the target number of rounds, and the aforementioned first set of operation action information as the target set of operation action information. The preset number of rounds represents the number of rounds in which the target error threshold condition is updated. The preset number of rounds can be 0.
[0053] Third, based on the target error threshold condition and the target operation action information set, perform the following modification steps: The first sub-step involves determining the action judgment result information of each target operation action information in the target operation action information set as the target judgment result, based on the aforementioned target error threshold condition. It should be noted that the method for determining whether the hand position included in the target operation action information is correct based on the aforementioned target error threshold condition is the same as the method for determining whether the hand position included in the operation action information is correct based on the preset error threshold condition. The method for determining whether the gesture included in the target operation action information is correct is the same as the method for determining whether the gesture included in the operation action information is correct, and will not be elaborated here; please refer to step 303. In practice, in response to determining that the gesture recognition information included in the aforementioned target operation action information meets the aforementioned preset standard condition, and that the hand position information included in the aforementioned target operation action information meets the aforementioned target error threshold condition, the processing terminal can correctly determine the aforementioned target operation action information as the target judgment result. The correctness of the aforementioned target operation action information indicates that the aforementioned target operation action information is correct.
[0054] The second sub-step, in response to determining that all obtained target judgment results are correct, updates the aforementioned target error threshold conditions and target rounds, and acquires a second set of operation action information collected from the head-mounted display device. The second operation action information in this set represents the user's gestures and hand positions during practical testing. The number of second operation action information in this set is the same as the number of first operation action information in the first set. In practice, firstly, the processing unit can determine the updated target round by summing the target round with 1. Then, it acquires the target error threshold conditions corresponding to the target rounds from the preset threshold correspondence information set as the updated target error threshold conditions. The preset threshold correspondence information in this set represents the correspondence between the target rounds and the target error threshold conditions. For example, the preset threshold correspondence information could be: target round 2, target error threshold condition is that the distance between the hand position and the preset position included in the changed assessment standard for the corresponding knowledge point in the above-mentioned detection standard editing information is in the range of -2cm to +2cm.
[0055] The third sub-step involves using the updated target error threshold condition as the target error threshold condition, the updated target round as the target round, and the second set of operation action information as the target operation action information set, and then executing the above change steps again.
[0056] The fourth sub-step, in response to the determination that there are erroneous target judgment results among the obtained target judgment results, generates target practical operation test result information based on the aforementioned target rounds and target judgment results. This target practical operation test result information can represent the final assessment result. In practice, firstly, the processing end can determine the preset detection level information corresponding to the aforementioned target rounds from the preset detection level information set. The preset detection level information set in the aforementioned preset detection level information set can represent the correspondence between target rounds and practical operation levels. For example, the target round is 2, and the practical operation level: the user assessment level is senior maintenance personnel. Then, in response to the determination that the accuracy rate of the aforementioned target judgment results is greater than the preset accuracy rate, a first preset score is determined as the test score. In response to the determination that the accuracy rate of the aforementioned target judgment results is less than or equal to the aforementioned preset accuracy rate, a second preset score is determined as the test score. The aforementioned preset accuracy rate can be 80%. The aforementioned first preset score can represent 80 points. The aforementioned second preset score can be 60 points. Finally, the obtained test score and practical operation level are determined as the target practical operation test result information.
[0057] The above-described technical solution, as an inventive point of this disclosure, solves technical problem two: "Low evaluation accuracy and insufficient discrimination lead to the system needing to perform more tests and longer time to obtain the final practical test results, thereby reducing the processing efficiency and resource utilization of practical test result information." The reasons for low evaluation accuracy and insufficient discrimination, thus increasing system processing time and resource consumption, are as follows: using a fixed fault tolerance threshold results in low evaluation accuracy and insufficient discrimination, which in turn increases the amount of invalid assessment data processed, requiring the system to perform more tests and longer time to obtain the final practical test results, thereby reducing the processing efficiency and resource utilization of practical test result information. If the above factors are resolved, evaluation accuracy and discrimination can be improved, thereby reducing system processing time and resource consumption. To achieve this effect, the practical test result information generation system of this disclosure dynamically updates the target error threshold condition in response to determining that the target judgment result of each target operation action information in the above-described target operation action information set is correct. Therefore, evaluation accuracy and discrimination can be improved, thereby reducing system processing time and resource consumption.
[0058] In addressing the technical problems mentioned above, and specifically for scenario three: industrial machine tool maintenance often presents the following challenges: individual maintenance steps cannot be tested in isolation, hindering automated verification of the overall logic and timing of the maintenance process. This results in users performing correct actions but failing tests due to incorrect procedures, reducing the accuracy and reliability of step detection. The system then needs to re-obtain data from the user's simulated operation to re-evaluate its correctness, increasing system downtime and resource consumption. Considering the specific requirements of this application scenario—adapting to tightly logical operation processes and stringent maintenance timing requirements—we have decided to adopt the following solution: Optionally, after step 305, the aforementioned processing terminal can also be configured to perform the following steps: The first step is to acquire standard operating procedure information and receive a set of operational action information. The standard operating procedure information can represent the standard state sequence of the virtual teaching aid during the assessment. The virtual teaching aid can represent each teaching aid included in the map information. The standard state in the standard state sequence represents the state of the virtual teaching aid after the target user correctly operates it during the assessment. For example, the standard state sequence could be "Standard state (power off) - Standard state (multimeter shows voltage) - Standard state (capacitor disconnected) - Standard state (new component soldered) - Standard state (power-on test light on)". In practice, firstly, the processing end can acquire the standard operating procedure information from the creation end. Then, it receives the set of operational action information from the head-mounted display device.
[0059] The second step involves determining the process scoring logic information and the process score information based on the aforementioned standard operating procedure information. The process scoring logic information represents the score corresponding to each standard state in the aforementioned standard state sequence. The process score information represents the score for the compliance of the process corresponding to the aforementioned standard operating procedure information. The process compliance indicates that the order of operations in the set of operation action information received by the head-mounted display device is the same as the order of the standard states in the aforementioned standard state sequence. The initial value of the process score information can be 0. In practice, firstly, the processing end can obtain the number of standard states included in the aforementioned standard operating procedure information through a state machine model. For example, the standard operating procedure information could be "Power off - Multimeter shows voltage - Disconnect capacitor - Solder new component - Power-on test light turns on - Complete," corresponding to 5 standard states. Then, the ratio of 100 to the number of standard states is determined as the process scoring logic information; therefore, the score for each standard state is 20 points. Each operation action information in the aforementioned operation action information set corresponds to a timestamp. The operation action information in the above set is sorted according to the corresponding timestamp.
[0060] The third step involves determining the initial standard states from the aforementioned standard operating procedure information as the target state sequence, and determining the aforementioned process score information as the target process score information. The aforementioned initial standard states can characterize the initial state of the aforementioned virtual teaching aid. For example, the initial standard state could be "power off".
[0061] Fourth, based on the operation action information, target state sequence, and target process score information that meet the preset selection conditions in the above operation action information set, perform the following verification steps: The first sub-step involves determining the state change data information of the virtual teaching aid based on the aforementioned operation action information. This state change data information characterizes the updated state of the virtual teaching aid. In practice, the processing terminal can obtain the state change data information of the virtual teaching aid corresponding to the aforementioned operation action information from a preset operation correspondence information set. The preset operation correspondence information in the preset operation correspondence information set characterizes the correspondence between the aforementioned operation action information and the state change data information of the virtual teaching aid. For example, the preset operation correspondence information could be: the aforementioned operation action information is "gesture recognition information is picking up the multimeter, hand position information is the position of the gesture when the probe touches point A," and the state change data information could be "the multimeter displays voltage."
[0062] The second sub-step involves generating an updated target state sequence based on the aforementioned state change data and target state sequence. This updated target state sequence characterizes the real-time state sequence comprised of the target state sequence and the aforementioned state change data. In practice, the processing unit can determine the target state sequence and the aforementioned state change data as the real-time state sequence. For example, the real-time state sequence could be "Power off - Multimeter shows voltage".
[0063] The third sub-step involves determining the updated target process score information based on the updated target state sequence and the standard operation procedure information. In practice, firstly, the preceding standard state in the real-time state sequence is determined as the first state. Then, the preceding standard state in the standard state sequence corresponding to the aforementioned state change data is determined as the second state. Finally, in response to determining that the first state and the second state are the same, the processing terminal can determine the sum of the target process score information and the process score logic information as the updated target process score information. In response to determining that the first state and the second state are different, the processing terminal can determine the target process score information as the updated target process score information.
[0064] The fourth sub-step involves, in response to determining that the aforementioned set of operation action information contains operation action information that satisfies the aforementioned preset selection conditions, using the updated target state sequence as the target state sequence and the updated target process score information as the target process score information, and then executing the aforementioned verification step again based on the operation action information in the aforementioned set of operation action information that satisfies the preset selection conditions. The aforementioned preset selection conditions can be that the timestamp of the selected operation action information is less than the timestamp of each unselected operation action information in the aforementioned set of operation action information.
[0065] Fifth, based on the target process score information and the practical operation test results information, generate modified practical operation test results information. This modified practical operation test results information represents the user's score at the end of the assessment. In practice, the processing terminal can input the target process score information and the practical operation test results information into the modified practical operation test results function to obtain the modified practical operation test results information.
[0066] As an example, the function for detecting the results of the above-mentioned change operation can be: .
[0067] in, It can characterize changes in practical test results. It can characterize the information of practical test results. It can characterize the weights corresponding to the above practical test results. It can represent the target process score information. This can represent the weights corresponding to the target process score information mentioned above. Here, for the above... , The specific value is not limited, for example, It can be 0.5. It can be 0.5.
[0068] The first to fifth steps of the above technical solution serve as an inventive point of this disclosure, solving the third technical problem mentioned above: "reduced accuracy and reliability of assessment, and the need to re-determine the correctness of the operation process, leading to increased system operation time and resource consumption." Factors causing reduced accuracy and reliability of assessment and increased system operation time and resource consumption often include: individual testing of each maintenance step, without the ability to automatically verify the logic and timing of the entire maintenance operation process. This results in users performing correct actions but passing the test due to incorrect procedures, reducing the accuracy and reliability of operation step detection. The system then needs to re-obtain data from the user's simulated operation process to re-determine the correctness of the operation process, leading to increased system operation time and resource consumption. To achieve this effect, this disclosure introduces a method for generating practical operation test result information. First, by comparing the standard operation process information with the updated target state sequence, target process score information is obtained. This target process score information includes the score for the operation sequence. Then, the target process score information and the practical operation test result information are weighted and fused with preset weights to obtain the final modified practical operation test result information. Therefore, a comprehensive score can be obtained for the corresponding operation sequence and steps. This can improve the accuracy and reliability of the assessment, thereby reducing the need for re-evaluating the correctness of the operation process, which would increase the system's operation time and resource consumption.
[0069] Optionally, after step 305, the aforementioned head-mounted display device can also be configured to determine process log information based on a text information set corresponding to the voice information set, a coordinate trajectory corresponding to the hand position information, and a timestamp. The voice information in the voice information set and the text information in the text information set can have a one-to-one correspondence. The text information in the text information set can represent the text of the corresponding voice information. The coordinate trajectory can represent a set of coordinate points formed by the change of the hand position over time when moving from the position represented by the hand position information as the starting or ending point. In practice, firstly, the executing entity can convert each voice information in the voice information set into text using ASR technology to obtain a text information set corresponding to the voice information set. Then, the coordinate trajectory corresponding to the hand position information is obtained through a mutually communicating motion controller. The motion controller can be a Leap Motion. Next, the timestamp corresponding to the hand position information is obtained through a built-in RTC clock synchronization mechanism. Finally, the text information set, the coordinate trajectory corresponding to the hand position information, and the timestamp are determined as process log information. Therefore, process log information can be obtained by recording the text information set corresponding to the user's voice information set, the coordinate trajectory of the corresponding hand position information, and the timestamp during the assessment process. This eliminates the need to first collect questionnaires for each knowledge point and then determine the rationality of each knowledge point. Instead, the rationality of each knowledge point can be analyzed using the process log information (e.g., if the accuracy rate of a knowledge point is below 40%, the knowledge point is difficult and the allocated processing time is short, thus the rationality of the corresponding knowledge point is unreasonable). There are no specific limitations on the rationality of each knowledge point; it can be set according to specific needs. For example, the rationality of each knowledge point can indicate that the knowledge point is difficult, thus requiring a longer allocated processing time. This reduces the time spent determining the rationality of each knowledge point in the knowledge point set and decreases the cost of collecting questionnaires.
[0070] Optionally, after step 305, the head-mounted display device can also be configured to upload the process log information to the processing terminal.
[0071] The above-described embodiments of this disclosure have the following beneficial effects: The practical operation test result information generation system of some embodiments of this disclosure can improve the effectiveness of user operation skill testing, increase the accuracy of practical operation test result information, and reduce the time and computational resources required to generate practical operation test result information. Specifically, the reasons for the low validity of the assessment results and the increased time and computational resources required to generate practical operation test result information are as follows: the interaction logic relying on general input devices such as mice and keyboards differs from professional skills such as hand operations and tool usage in real-world environments. It cannot capture and detect key skill indicators such as gesture standardization, force control, and spatial trajectory accuracy involved in real-world operations, resulting in a disconnect between the detected content and the practical scenario, leading to poor effectiveness in testing user operation skills. Furthermore, the system can only record simple interaction events (such as click coordinates and operation sequences), resulting in low accuracy of the obtained practical operation test result information. This necessitates additional video recording and review of the videos, further increasing the time and computational resources required to generate practical operation test result information. Based on this, the practical operation test result information generation system of some embodiments of this disclosure includes a creation terminal, a head-mounted display device, and a processing terminal. The creation terminal is used to obtain map information of the current practical operation test scenario from the processing terminal. This allows for the acquisition of various teaching aids and their environments corresponding to the current practical operation test scenario. The creation terminal is used to edit the practical operation test standard information of the current practical operation test scenario based on the map information, obtaining test standard editing information. This allows for the acquisition of assessment standards corresponding to each test point in the current practical operation test scenario. The head-mounted display device is used to generate an action judgment result information set based on the above-mentioned test standard editing information and the collected set of operation action information. This allows for the acquisition of a judgment result for each operation action information in the operation action information set. The processing terminal is used to generate a speech recognition result information set based on the collected set of speech information and the above-mentioned test standard editing information. This allows for the acquisition of a recognition result for each speech information in the speech information set. The processing terminal is used to generate practical operation test result information based on the above-mentioned action judgment result information set and the above-mentioned speech recognition result information set. This allows for the acquisition of the final score after the assessment. Because it doesn't rely on common input devices like mice and keyboards for interaction when generating practical operation test results, but instead simulates practical operation tests in a real-world environment within augmented reality, it can effectively capture and test key skill indicators such as gesture standardization, force control, and spatial trajectory accuracy involved in real-world operations, thus improving the effectiveness of the test results.Furthermore, because the system doesn't just record simple interaction events (such as click coordinates or operation sequences), but also synchronously collects voice information and spatial positioning data provided by SLAM, it can improve the accuracy of practical operation test results. This eliminates the need for additional video recording and review, reducing the time and computational resources required to generate practical operation test results. Consequently, it can improve the validity of the assessment results, increase the accuracy of practical operation test results, and reduce the time and computational resources required to generate practical operation test results.
[0072] Further reference Figure 4 The document illustrates a flowchart 400 of some embodiments of a method for generating practical test result information for a head-mounted display device. The flowchart 400 of this practical test result information generation method includes the following steps: Step 401: Collect the voice information set and the operation action information set.
[0073] In some embodiments, the head-mounted display device described above can collect a set of voice information and a set of operation action information. In practice, the head-mounted display device can collect the user's voice information through a built-in microphone array. The user can be a user wearing the head-mounted display device. Here, the method for collecting the operation action information set is the same as the method for determining the operation action information set in the above-described practical operation test result information generation system, and will not be described again.
[0074] Step 402: Based on the detection standard, edit the information and operation action information set to generate the action judgment result information set.
[0075] In some embodiments, the specific implementation of step 402 and its resulting technical effects can be found in [reference needed]. Figure 3 Step 303 in the corresponding embodiments will not be repeated here.
[0076] Step 403: Send the voice information set and the action judgment result information set to the processing terminal included in the practical test result information generation system.
[0077] In some embodiments, the executing entity may send the voice information set and the action determination result information set to the processing terminal included in the practical operation test result information generation system. In practice, the executing entity may send the voice information set and the action determination result information set to the processing terminal included in the practical operation test result information generation system via the IP address of the processing terminal.
[0078] like Figure 5As shown, the head-mounted display device 500 includes a processing unit (CPU) 501, a memory (ROM) 502, an input unit 503, and an output unit 504, wherein the processing unit 501, memory 502, input unit 503, and output unit 504 are interconnected via a bus 505. Here, the methods according to some embodiments of this disclosure can be implemented as a computer program and stored in the memory 502. The processing unit 501 in the head-mounted display device 500 implements the virtual keyboard display function defined in some embodiments of this disclosure by calling the aforementioned computer program stored in the memory 502. In some implementations, the input unit 503 may include devices such as a camera, microphone, gyroscope, accelerometer, and magnetometer, and the output unit 504 may be devices that can display content, such as optomechanical and optical elements. The aforementioned optomechanical and optical elements may be microdisplays. Thus, when the processing unit 501 calls the aforementioned computer program to execute the split-screen display function, it can control the input unit 503 to acquire user gestures, voice, and other operation commands, and control the output unit 504 to display the content.
[0079] It should be noted that, in some embodiments of this disclosure, the computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0080] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0081] The aforementioned computer-readable medium may be included in the aforementioned head-mounted display device; or it may exist independently and not assembled into the head-mounted display device. The aforementioned computer-readable medium carries one or more programs, which, when executed by the head-mounted display device, cause the head-mounted display device to: acquire a set of voice information and a set of operation action information; edit information according to the aforementioned detection standards and the aforementioned set of operation action information to generate a set of action judgment result information; and send the aforementioned set of voice information and the aforementioned set of action judgment result information to the processing terminal included in the aforementioned practical operation detection result information generation system.
[0082] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0083] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0084] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0085] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.
Claims
1. A practical test result information generation system, wherein, The practical test result information generation system includes: an authoring terminal, a processing terminal, and a head-mounted display device that are interconnected. The creation terminal is used to obtain map information of the current practical testing scenario from the processing terminal; based on the map information, the practical testing standard information of the current practical testing scenario is edited to obtain testing standard editing information; The head-mounted display device is used to generate a set of action judgment result information based on the detection standard, the edited information, and the collected set of operation action information. The processing terminal is used to generate a speech recognition result information set based on the collected speech information set and the detection standard editing information; and to generate practical detection result information based on the action judgment result information set and the speech recognition result information set.
2. The system according to claim 1, wherein, The practical test result information generation system also includes a map acquisition terminal; and The map acquisition terminal is configured to generate map information corresponding to each practical testing teaching aid based on the scanning result information of each practical testing teaching aid included in the current practical testing scenario. The map information is uploaded to the processing terminal.
3. The system according to claim 1, wherein, The head-mounted display device is configured to: Based on each operation action information in the operation action information set, the following determination steps are performed: In response to determining that the operation action information meets the preset judgment conditions, the action detection information is correctly determined as the action judgment result information; In response to determining that the operation action information does not meet the preset judgment condition, the action detection information is incorrectly determined as the action judgment result information; The obtained action judgment results are defined as a set of action judgment results.
4. The system according to claim 3, wherein, The preset judgment condition is that the gesture recognition information included in the operation action information meets the preset standard condition, and the hand position information included in the operation action information meets the preset error threshold condition.
5. The system according to claim 1, wherein, The head-mounted display device is configured to: Based on the text information set corresponding to the voice information set, the coordinate trajectory of the corresponding hand position information, and the timestamp, the process log information is determined; The process log information is uploaded to the processing terminal.
6. The system according to claim 1, wherein, The processing terminal is configured as follows: For each piece of speech information in the speech information set, perform the following recognition steps: Edit the information according to the detection criteria, perform matching processing on the voice information, and obtain matching degree information; In response to determining that the matching degree information meets the preset threshold condition, the voice information is correctly identified as voice recognition result information; In response to determining that the matching degree information does not meet the preset threshold condition, the voice information is incorrectly identified as voice recognition result information; The obtained speech recognition results are defined as a speech recognition result information set.
7. The system according to claim 6, wherein, The preset threshold condition is that the matching degree information is greater than or equal to a preset value.
8. A method for generating practical test result information, applied to a head-mounted display device included in the practical test result information generation system according to any one of claims 1-7, comprising: Collect sets of voice information and sets of user action information; Based on the detection standard editing information and the operation action information set, generate an action judgment result information set; The set of voice information and the set of action determination results information are sent to the processing terminal included in the practical operation detection result information generation system.
9. A head-mounted display device, comprising: One or more processors; Storage device for storing one or more programs; Optical mechanisms and optical components used to create images in front of the user's eyes; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 8.
10. A computer-readable medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the method as described in claim 8.