VR teaching method, VR teaching device and VR teaching system
By obtaining user's personalized needs information, building a scene content generation model, generating and dividing virtual teaching scenarios, and adding interactive nodes, the problem that existing VR teaching methods cannot meet personalized needs is solved, and an efficient and customized virtual teaching experience is achieved.
Patent Information
- Application Number
- CN202510294064.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-27
AI Technical Summary
The existing VR teaching methods cannot meet the personalized needs of users, and the traditional VR teaching scenario design is fixed, so it cannot quickly adapt to the update of teaching content and changes in learners' needs.
By obtaining the target instruction information and preference information of the target user, a scene content generation model is built, a virtual scene that meets user needs is generated, and the scene is divided, interactive nodes are added to form a personalized virtual teaching scenario.
It realizes the generation of customized virtual teaching scenarios based on user personalized needs, improves learning experience and teaching effect, and solves the problem that traditional VR teaching methods cannot meet diversified needs.
Smart Images

Figure CN120215942A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of virtual reality, and in particular, to a VR teaching method, device, computer-readable storage medium, and VR teaching system. Background Art
[0002] In the development and application of VR teaching methods, there are already some systems that can generate virtual teaching scenarios according to teaching syllabuses and learning objectives. However, these systems often ignore the personalized needs of learners. For example, different learners may have their own preferences for the visual performance, auditory environment, or interaction methods of the teaching scenario. Traditional VR teaching methods often generate scenarios according to fixed templates, unable to meet diverse needs, resulting in a reduction in the learning experience and limited teaching effects. In addition, the construction of teaching scenarios usually requires professional design and development teams, which is time-consuming and laborious, and it is difficult to quickly adapt to the update of teaching content and the changes in learners' needs. Traditional VR teaching scenarios are often designed once and are difficult to modify or adjust once generated, which limits the flexibility and customization of teaching scenarios and is difficult to meet the needs of learners of different ages, learning abilities, and interests.
[0003] Regarding the content division of teaching scenarios, traditional VR teaching methods often adopt the methods of linear playback or single-scene display, lacking the ability to intelligently divide and dynamically display according to the logical relationship or time sequence of the scenario content. This makes the presentation mode of teaching content single, and it is difficult for learners to follow the time clue or logical chain for in-depth learning and understanding. Summary of the Invention
[0004] The main purpose of this application is to provide a VR teaching method, device, computer-readable storage medium, and VR teaching system to at least solve the problem that the VR system in the prior art cannot meet the personalized needs of users.
[0005] To achieve the above object, according to one aspect of the present application, a VR teaching method is provided, including: obtaining target instruction information and preference information of a target user, where the target instruction information is query information sent by the target user to the VR system, and the preference information includes visual preference, auditory preference, and interaction mode preference; constructing a scene content generation model, which is used to generate a virtual scene of the VR system according to the target instruction information and the preference information, where the scene content generation model is trained using multiple sets of training data, and each set of training data in the multiple sets of training data includes historical target instruction information and historical preference information; inputting the target instruction information and the preference information into the scene content generation model to obtain a target virtual scene, so that the target virtual scene meets the requirements of the preference information of the target user; dividing the target virtual scene according to the type of the target virtual scene to obtain multiple virtual sub-scenes, and adding interaction nodes to each of the virtual sub-scenes, where the interaction nodes are used to answer the target instruction information related to the corresponding virtual sub-scene or jump to the virtual sub-scene corresponding to the target instruction information not related to the corresponding virtual sub-scene; using the target virtual scene after adding the interaction nodes as the final virtual scene, and playing the final virtual scene to teach the target user.
[0006] Optionally, constructing the scene content generation model includes: obtaining multiple pieces of the historical target instruction information and multiple pieces of the historical preference information; determining the audience labels of each piece of the historical preference information, where the audience labels include education practitioner labels and ordinary user labels; extracting feature vectors from each piece of the historical target instruction information and each piece of the historical preference information with the audience labels to obtain multiple historical target instruction vectors and multiple historical preference vectors; training an LSTM neural network model using each of the historical target instruction vectors and each of the historical preference vectors to obtain the scene content generation model.
[0007] Optionally, dividing the target virtual scene according to the type of the target virtual scene to obtain multiple virtual sub-scenes includes: constructing a scene category model, where the scene category model is trained using multiple sets of training data, and each set of training data in the multiple sets of training data includes historical target virtual scenes with type labels obtained within a historical time period, and the type labels include time type labels, logical type labels, and multi-element type labels; inputting the target virtual scene into the scene category model to obtain the type label corresponding to the target virtual scene; determining the division method of the target virtual scene according to the type label, and dividing the target virtual scene according to the division method to obtain multiple virtual sub-scenes.
[0008] Optionally, determine the partitioning method of the target virtual scene according to the type tag, and partition the target virtual scene according to the partitioning method to obtain a plurality of the virtual sub-scenes, including: when the type tag corresponding to the target virtual scene is the time type tag, use a time expression recognition algorithm to extract a first time point and a second time point of the target virtual scene, where the first time point is all the time points in the target virtual scene, and the second time point is an important event node of the target virtual scene; arrange the first time point and the second time point in chronological order to obtain a time axis; use the second time point as a first segmentation point to segment the time axis to obtain a plurality of time periods; determine the target virtual scene corresponding to each time period as a plurality of the virtual sub-scenes corresponding to the target virtual scene.
[0009] Optionally, determine the partitioning method of the target virtual scene according to the type tag, and partition the target virtual scene according to the partitioning method to obtain a plurality of the virtual sub-scenes, further including: when the type tag corresponding to the target virtual scene is the logic type tag, obtain all the text contents of the target virtual scene; identify the logical relation words in the text contents, where the logical relation words include causal relation words, progressive relation words, and conditional relation words; use the logical relation words as a second segmentation point to segment the text contents to obtain a plurality of text segments; determine the target virtual scene corresponding to each text segment as a plurality of the virtual sub-scenes corresponding to the target virtual scene.
[0010] Optionally, determine the partitioning method of the target virtual scene according to the type tag, and partition the target virtual scene according to the partitioning method to obtain a plurality of the virtual sub-scenes, further including: when the type tag corresponding to the target virtual scene is the multi-element type tag, obtain all the target elements in the target virtual scene, where the target elements include people, events, locations, items, and concepts; determine the importance degree of the target elements in the target virtual scene to obtain the importance weight corresponding to each target element; based on a preset value, combine each target element to obtain a target element group, where the sum of the importance weights corresponding to each target element in the target element group is less than the preset value, and the target element group includes at least one person, one event, one location, one item, and one concept; determine the target virtual scene corresponding to each target element group as a plurality of the virtual sub-scenes corresponding to the target virtual scene.
[0011] Optionally, after obtaining the final virtual scene, the method further includes: when the target user uses a VR device to view the final virtual scene, controlling the VR system to stop playing at the time point where the interaction node is located, so that the target user can use the interaction node for interaction; when the target user does not use the interaction node for interaction within a predetermined time, controlling the VR system to automatically play the final virtual scene.
[0012] To achieve the above object, according to one aspect of the present application, there is provided a VR teaching device, including: an acquisition unit, configured to acquire the target instruction information and preference information of a target user, where the target instruction information is query information sent by the target user to the VR system, and the preference information includes visual preference, auditory preference, and interaction mode preference; a construction unit, configured to construct a scene content generation model, where the scene content generation model is used to generate a virtual scene of the VR system according to the target instruction information and the preference information, and wherein the scene content generation model is obtained by training using multiple sets of training data, and each set of training data in the multiple sets of training data includes historical target instruction information and historical preference information; a training unit, configured to input the target instruction information and the preference information into the scene content generation model to obtain a target virtual scene, so that the target virtual scene meets the requirements of the preference information of the target user; a division unit, configured to divide the target virtual scene according to the type of the target virtual scene to obtain multiple virtual sub-scenes, and add interaction nodes to each of the virtual sub-scenes, where the interaction nodes are used to answer the target instruction information related to the corresponding virtual sub-scene or jump to the virtual sub-scene corresponding to the target instruction information not related to the corresponding virtual sub-scene; a control unit, configured to use the target virtual scene after adding the interaction nodes as the final virtual scene, and play the final virtual scene to teach the target user.
[0013] According to another aspect of the present application, there is provided a computer-readable storage medium, where the computer-readable storage medium includes a stored program, and wherein when the program runs, it controls any one of the methods in the device where the computer-readable storage medium is located.
[0014] According to another aspect of the present application, there is provided a VR teaching system, including: one or more processors, a memory, and one or more programs, where the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include those for executing any one of the methods.
[0015] Applying the technical solution of the present application, in the above-mentioned VR teaching method, it includes: obtaining the target instruction information and preference information of the target user, where the target instruction information is the query information sent by the target user to the VR system, and the preference information includes visual preference, auditory preference, and interaction mode preference; constructing a scene content generation model, where the scene content generation model is used to generate the virtual scene of the VR system according to the target instruction information and the preference information, and among them, the scene content generation model is trained using multiple groups of training data, and each group of training data in the multiple groups of training data includes historical target instruction information and historical preference information; inputting the target instruction information and the preference information into the scene content generation model to obtain a target virtual scene, so that the target virtual scene meets the requirements of the preference information of the target user; dividing the target virtual scene according to the type of the target virtual scene to obtain multiple virtual sub-scenes, and adding interaction nodes to each of the virtual sub-scenes, where the interaction nodes are used to answer the target instruction information related to the corresponding virtual sub-scene or jump to the virtual sub-scene corresponding to the target instruction information not related to the corresponding virtual sub-scene; using the target virtual scene after adding the interaction nodes as the final virtual scene, and playing the final virtual scene to teach the target user. The present application generates a target virtual scene that meets the needs of the target user according to the target instruction information and preference information of the target user, divides the target virtual scene to obtain multiple virtual sub-scenes, and sets interaction nodes for each virtual sub-scene, so that the user can obtain the target instruction information of the corresponding virtual sub-scene or jump to the virtual sub-scene corresponding to the target instruction information not related to the corresponding virtual sub-scene through the interaction nodes, and the target virtual scene after adding the interaction nodes is the final virtual scene. By generating a personalized target virtual scene and interaction nodes according to the target instruction information and preference information of the target user, the problem that the content and interaction mode of the virtual scene are both fixed modes is avoided, the user can obtain a better VR experience, and the problem that the VR system in the prior art cannot meet the personalized needs of users is solved. Description of the Drawings
[0016] Figure 1 Shows a hardware structure block diagram of a mobile terminal for implementing a VR teaching method provided in an embodiment of the present application;
[0017] Figure 2 Shows a flowchart of a VR teaching method provided in an embodiment of the present application;
[0018] Figure 3 Shows a structure block diagram of a VR teaching device provided in an embodiment of the present application.
[0019] Among them, the above-mentioned drawings include the following reference numerals:
[0020] 102, processor; 104, memory; 106, transmission device; 108, input / output device. Detailed implementation manners
[0021] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the drawings and in conjunction with the embodiments.
[0022] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present application.
[0023] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so as to describe the embodiments of the present application here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0024] As introduced in the background art, traditional VR teaching methods in the prior art often generate scenes according to a fixed template and cannot meet diverse needs. To solve this technical problem, embodiments of the present application provide a VR teaching method, device, computer-readable storage medium, and VR teaching system.
[0025] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention.
[0026] The method embodiments provided in the embodiments of the present application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 is a hardware structure block diagram of a mobile terminal of a VR teaching method according to an embodiment of the present invention. As Figure 1 shown, the mobile terminal may include one or more ( Figure 1Only one processor 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a field-programmable gate array FPGA) and a memory 104 for storing data are shown. Among them, the above mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 The structure shown is only illustrative and does not limit the structure of the above mobile terminal. For example, the mobile terminal may further include more or fewer components than Figure 1 shown in, or have a different configuration from Figure 1 shown.
[0027] The memory 104 can be used to store computer programs. For example, software programs and modules of application software, such as the computer program corresponding to a VR teaching method in an embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the mobile terminal through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof. The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the mobile terminal. In one instance, the transmission device 106 includes a network adapter (abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0028] In this embodiment, a VR teaching method running on a mobile terminal, a computer terminal, or a similar computing device is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0029] Figure 2 is a flowchart of a VR teaching method according to an embodiment of the present application. As Figure 2 shown, the method includes the following steps:
[0030] Step S201: Obtain the target instruction information and preference information of the target user. The above-mentioned target instruction information is the query information sent by the above-mentioned target user to the VR system, and the above-mentioned preference information includes visual preference, auditory preference, and interaction mode preference;
[0031] Specifically, first, receive the query information input by the target user through the VR interface or external device, that is, the target instruction information, which covers the knowledge topics and content details that the user expects to explore or master in depth. At the same time, receive the preference information of the target user. After the user wears the VR device for the first time, the system collects the preference information through the VR device, including visual preferences, such as preferences for realistic style, cartoon style, or abstract expression, as well as personal tendencies towards color and brightness; auditory preferences, involving the choice of background music, the intonation and rhythm preferences of voice explanations; and interaction mode preferences. The user may be more inclined to use a handle, voice commands, or gesture control to communicate with the virtual environment. When the user uses VR multiple times, the system automatically collects the user's historical data, including the browsing history, search records, preference information in the VR system, and data in the historical virtual reality scenarios of the user. When the user uses the VR device again, through the unique identifier of the VR device, such as ID, ID card information, mobile phone number, etc., automatically obtain the user's personal preference information, historical browsing records, search records, and other information.
[0032] Step S202: Construct a scene content generation model. The above-mentioned scene content generation model is used to generate the virtual scene of the above-mentioned VR system according to the above-mentioned target instruction information and the above-mentioned preference information. Among them, the above-mentioned scene content generation model is trained using multiple groups of training data, and each group of training data in the above-mentioned multiple groups of training data includes historical target instruction information and historical preference information;
[0033] Specifically, by constructing a scene content generation model, the core function of this model is to intelligently analyze and integrate the target instruction information and preference information of the target user, so as to generate a highly customized VR virtual scene. The training process of this model uses a large number of historical data sets. Each data set contains the query information sent by past users (historical target instruction information) and the personal preferences of these users in terms of vision, hearing, and interaction mode (historical preference information). Through deep learning technology, the model can extract the correlation patterns between user instructions and preferences, and how these patterns affect the content design and presentation of virtual scenes. Finally, the scene content generation model can generate personalized scenes that meet the requirements according to the target instruction information and preference information of the user.
[0034] Step S203: Input the above-mentioned target instruction information and the above-mentioned preference information into the above-mentioned scenario content generation model to obtain a target virtual scenario, so that the target virtual scenario meets the requirements of the above-mentioned preference information of the above-mentioned target user;
[0035] Specifically, by combining the query information proposed by the target user, that is, the target instruction information, with the user's personalized preference information, a virtual learning scenario that meets the specific needs of the user is intelligently generated. Specifically, the collected target instruction information and preference information are used as input parameters and transmitted to the trained scenario content generation model. The model uses the patterns it has learned to analyze the user's intention, and at the same time considers the user's preferences in terms of vision, hearing, and interaction methods, and accurately generates a target virtual scenario that highly matches the user's needs. Whether the user pursues a visually rich and vivid visual effect, or prefers an immersive auditory experience, or has a preference for a specific interaction method, the model can generate a virtual scenario that meets these preferences, ensuring that each user can obtain the best learning experience.
[0036] Step S204: Divide the above-mentioned target virtual scenario according to the type of the above-mentioned target virtual scenario to obtain multiple virtual sub-scenarios, and add interaction nodes to each of the above-mentioned virtual sub-scenarios. The above-mentioned interaction nodes are used to answer the above-mentioned target instruction information related to the corresponding above-mentioned virtual sub-scenario or jump to the above-mentioned virtual sub-scenario corresponding to the above-mentioned target instruction information that is not related to the corresponding above-mentioned virtual sub-scenario;
[0037] Specifically, according to the type of the generated target virtual scenario, it is divided into multiple smaller and more manageable virtual sub-scenarios. This division process aims to identify the key themes or teaching units in the scenario, ensuring that each sub-scenario can focus on specific learning objectives and provide in-depth teaching content. More importantly, after obtaining the virtual sub-scenarios by division, interaction nodes will be embedded in each virtual sub-scenario. These nodes are designed to respond to the user's learning needs within the sub-scenario. They can not only answer the target instruction information directly related to the current sub-scenario, guiding the user to deeply understand the scenario content, but also provide a jump function, allowing the user to quickly access other virtual sub-scenarios that are not directly related to the theme of the current virtual sub-scenario but are consistent with the user's overall learning objectives, realizing cross-domain exploration of knowledge.
[0038] Step S205: Use the above-mentioned target virtual scenario after adding the above-mentioned interaction nodes as the final virtual scenario, and play the above-mentioned final virtual scenario to teach the above-mentioned target user.
[0039] Specifically, after adding interaction nodes to the generated target virtual scene, it is transformed into the final virtual teaching scene, and then presented to the target user to guide them to conduct immersive learning. This process is not simply playing the scene, but through the designed interaction nodes, closely integrating the user's learning path with the dynamic display of the virtual scene, creating an interactive learning experience.
[0040] Through this embodiment, in the above VR teaching method, the target instruction information and preference information of the target user are obtained. The above target instruction information is the query information sent by the above target user to the VR system. The above preference information includes visual preference, auditory preference, and interaction mode preference; a scene content generation model is constructed. The above scene content generation model is used to generate the virtual scene of the above VR system according to the above target instruction information and the above preference information. Among them, the above scene content generation model is trained using multiple groups of training data. Each group of training data in the above multiple groups of training data includes historical target instruction information and historical preference information; the above target instruction information and the above preference information are input into the above scene content generation model to obtain a target virtual scene, so that the above target virtual scene meets the requirements of the above preference information of the above target user; the above target virtual scene is divided according to the type of the above target virtual scene to obtain multiple virtual sub-scenes, and interaction nodes are added to each of the above virtual sub-scenes. The above interaction nodes are used to answer the above target instruction information related to the corresponding above virtual sub-scene or jump to the above virtual sub-scene corresponding to the above target instruction information that is not related to the corresponding above virtual sub-scene; the above target virtual scene after adding the above interaction nodes is used as the final virtual scene, and the above final virtual scene is played to teach the above target user. This application generates a target virtual scene that meets the needs of the target user according to the target instruction information and preference information of the target user, divides the target virtual scene to obtain multiple virtual sub-scenes, and sets interaction nodes for each virtual sub-scene, so that the user can obtain the target instruction information of the corresponding virtual sub-scene or jump to the virtual sub-scene corresponding to the target instruction information that is not related to the corresponding virtual sub-scene through the interaction nodes. The target virtual scene after adding the interaction nodes is the final virtual scene. By generating a personalized target virtual scene and interaction nodes according to the target instruction information and preference information of the target user, the problem that the content and interaction mode of the virtual scene are both in a fixed mode is avoided, enabling the user to obtain a better VR experience and solving the problem that the VR system in the prior art cannot meet the personalized needs of users.
[0041] In an alternative embodiment, in order to obtain the scene content generation model, the construction of the scene content generation model in step S202 includes:
[0042] Step S2021, obtain multiple pieces of the above-mentioned historical target instruction information and multiple pieces of the above-mentioned historical preference information;
[0043] Step S2022, determine the audience tags of each of the above-mentioned historical preference information, where the audience tags include education practitioner tags and ordinary user tags;
[0044] Specifically, collect the query information sent by multiple past users, that is, historical target instruction information, which covers the learning needs of users in specific knowledge fields; at the same time, also record multiple historical preference information, which reflects the personal preferences of users in terms of vision, hearing, and interaction methods. Then, we conduct in-depth mining on each set of historical preference information to determine the matching audience tags, including education practitioner tags and ordinary user tags. This step aims to identify the professional background and personal attributes of users, so as to more accurately understand their learning behaviors and preference patterns.
[0045] Step S2023, extract feature vectors from each of the above-mentioned historical target instruction information and each of the above-mentioned historical preference information with the above-mentioned audience tags to obtain multiple historical target instruction vectors and multiple historical preference vectors;
[0046] Step S2024, use each of the above-mentioned historical target instruction vectors and each of the above-mentioned historical preference vectors to train the LSTM neural network model to obtain a scene content generation model.
[0047] Specifically, a large number of collected historical target instruction information is subjected to feature vector extraction to extract feature vectors closely related to the user's query intention and scenario theme, namely historical target instruction vectors. At the same time, for historical preference information with clear audience labels (such as education practitioners, ordinary users, etc.), feature vector extraction is also performed to obtain historical preference vectors, which detail the specific preferences of users in terms of vision, hearing, and interaction methods. The specific extraction process for historical target instruction information is to clean the text in the historical target instruction information, remove stop words and special symbols, and then use word embedding technology to convert the processed text into a vector representation. Using a pre-trained word vector model, each word is mapped to a vector of a fixed dimension (such as 300 dimensions). The specific extraction process for historical preference information is to numerically encode each preference data. For example, the visual picture style can be divided into realistic (encoded as 0), cartoon (encoded as 1), abstract (encoded as 2), etc.; color preferences can be numerically encoded between 0 and 1 according to the hue range. For example, if warm colors are preferred, it is encoded as a higher value, and if cool colors are preferred, it is encoded as a lower value, etc. These encoded preference data are combined into a feature vector as another part of the input layer of the LSTM network. For example, after encoding the preference data, a 20-dimensional vector is formed. Subsequently, the system uses the extracted historical target instruction vectors and historical preference vectors to deeply train the LSTM (Long Short-Term Memory) neural network model. Through this training process, the model not only learns how to generate corresponding scenario content according to the user's instructions but also masters how to adjust the presentation method of the scenario according to the user's preferences, ensuring that the finally generated VR teaching scenario accurately reflects the user's needs and is highly customized to meet the user's personalized preferences.
[0048] The loss function of the above scenario content generation model is L = αL con + βL ass + γL aud ; where α, β, γ represent adjustable hyperparameters; L con represents the loss function in terms of content vision, x represents the feature vector of the audience's preference information, x i represents the i-th element of the vector x, s represents the feature vector of the generated scenario content, s i represents the i-th element of the vector s, and n is the dimension of the preference information feature vector x; L ass represents the loss function in terms of content hearing, y represents the association data matrix of different preferences, y ij represents the association strength between the i-th preference information and the j-th association data, and m is the number of columns of the association data matrix y of different preferences; L aud represents the loss function in terms of content interaction, z represents the attribute category vector of the current audience, zl Denoted as the l-th element in the vector z, ω is denoted as the weight vector, and k is denoted as the dimension of the attribute category data vector z of the current audience.
[0049] In a specific embodiment, in terms of structural design, the above LSTM neural network model adopts a multi-level combined design. The first layer of LSTM units receives the target instruction information and partial preference information (such as the basic information of visual and auditory preferences). Its main function is to initially extract the semantic features and emotional tendencies in the instructions and preferences. The number of hidden units in this layer can be set to 128. After being processed by this layer, an intermediate feature vector is output. The second layer of LSTM units receives the output of the first layer and the remaining preference information (such as interaction preference data), further integrating the information and mining deeper relationships. The number of hidden units in this layer is set to 256, which can perform a more accurate analysis of the combination possibilities of scene elements and the adaptability of interactions to the scene. For example, determine the appropriate interaction trigger areas and methods around specific scene elements according to the interaction preferences. The last layer of LSTM units generates the final scene content parameters according to the outputs of the previous two layers, including scene themes, elements, layouts, and interaction methods, etc. The number of hidden units in this layer is determined according to the types and complexities of the output parameters. For example, it is set to 512 to ensure that rich, diverse, and accurate scene content information can be generated.
[0050] In order to divide the target virtual sub-scene, in an alternative embodiment, the above target virtual scene is divided according to the type of the above target virtual scene to obtain multiple virtual sub-scenes. The above step S204 includes:
[0051] Step S2041, construct a scene category model. Among them, the above scene category model is trained using multiple sets of training data. Each set of the above multiple sets of training data includes historical target virtual scenes with type labels obtained within a historical time period. The above type labels include time type labels, logical type labels, and multi-element type labels;
[0052] Specifically, the construction of the scene category model is based on a large amount of historical target virtual scene data. Each set of training data is labeled with scene type labels, including time type labels, logical type labels, and multi-element type labels. These labels reflect the characteristics of the scene in terms of time series, logical structure, and element combination. By using multiple sets of such labeled historical data to deeply train the model, the system can learn the inherent features and generation patterns of different types of scenes, so as to be able to judge the type to which the target virtual scene belongs.
[0053] Step S2042, input the above target virtual scene into the above scene category model to obtain the above type label corresponding to the above target virtual scene;
[0054] Specifically, the initially generated target virtual scene is used as input and passed to a deeply trained scene classification model. Based on its learning of a large number of historical target virtual scenes, this model can quickly and accurately identify the core attributes of the target scene and assign corresponding type labels to it, such as time type, logical type, or multi-element type.
[0055] Step S2043: Determine the partitioning method of the above-mentioned target virtual scene according to the above type label, and partition the above-mentioned target virtual scene according to the above partitioning method to obtain multiple above-mentioned virtual sub-scenes.
[0056] Specifically, parse the type label output by the scene classification model, which reveals the structural attributes of the target virtual scene. Subsequently, based on these type labels, automatically select the most appropriate scene partitioning strategy to cut the complex target scene into a series of virtual sub-scenes that are easier to understand and interact with. This process takes into account the structural partitioning of the scene and ensures a high correlation for each virtual sub-scene.
[0057] In order to obtain virtual sub-scenes, in an optional implementation manner, determine the partitioning method of the above-mentioned target virtual scene according to the above type label, and partition the above-mentioned target virtual scene according to the above partitioning method to obtain multiple above-mentioned virtual sub-scenes. The above step S2043 includes:
[0058] Step S204301: When the type label corresponding to the above-mentioned target virtual scene is the above-mentioned time type label, use a time expression recognition algorithm to extract the first time point and the second time point of the above-mentioned target virtual scene. The first time point is all the time points in the above-mentioned target virtual scene, and the second time point is the important event node of the above-mentioned target virtual scene;
[0059] Specifically, use the time expression recognition algorithm to carefully sort out the time clues in the scene. This process first comprehensively scans the target virtual scene, accurately locates each time point, and forms a complete sequence of the first time points, which cover all the historical moments from the beginning to the end of the scene. Subsequently, the algorithm further conducts in-depth analysis and filters out the second time points with great significance or teaching value from the sequence of the first time points, that is, important event nodes, which are often associated with key teaching content.
[0060] Step S204302: Arrange the above first time point and the above second time point in chronological order to obtain a timeline;
[0061] Specifically, based on the first time point and the second time point, following the chronological order, a timeline is constructed to connect all historical moments and important events in the time context, that is, the timeline carries the time points including important events.
[0062] Step S204303, divide the above timeline using the above second time point as the first division point to obtain multiple time periods;
[0063] Specifically, based on the second time point, that is, those time nodes closely related to important events or teaching focuses, as the first division point. Based on these first division points, the complete timeline is accurately divided, thereby generating a series of time periods with clear key stages or teaching themes.
[0064] Step S204304, determine the above target virtual scenes corresponding to each time period as multiple above virtual sub - scenes corresponding to the above target virtual scene.
[0065] Specifically, regard each time period obtained by subdividing the timeline as an independent segment of the target virtual scene, and determine the virtual sub - scene corresponding to each time period as an independent virtual sub - scene.
[0066] In order to obtain virtual sub - scenes, in an optional implementation manner, determine the division method of the above target virtual scene according to the above type label, and divide the above target virtual scene according to the above division method to obtain multiple above virtual sub - scenes. The above step S2043 further includes:
[0067] Step S204305, when the above type label corresponding to the above target virtual scene is the above logical type label, obtain all the text contents of the above target virtual scene;
[0068] Specifically, deeply scan the target virtual scene to ensure that all text contents in the scene are captured. This process not only covers the statically displayed text information, but also can integrate the text materials that may be generated in dynamic interactions according to the actual scene to form a text content set, providing a data basis for subsequent scene segmentation.
[0069] Step S204306, identify the logical relationship words in the above text contents, and the above logical relationship words include causal relationship words, progressive relationship words, and conditional relationship words;
[0070] Specifically, natural language processing technology is used to extract logical relationship words in text content. These logical relationship words include but are not limited to causal relationship words (such as "therefore", "due to", "result"), progressive relationship words (such as "then", "afterwards", "further") and conditional relationship words (such as "if", "unless", "however"). They are like hidden threads of story narration, revealing the logical connection and advancement context between the various parts of the scene content. By locating these logical relationship words, it is possible to understand the narrative logic of the scene, distinguish the causal relationship between different elements, the continuous steps of event development, and the multiple paths triggered by conditions, thereby providing guidance information for subsequent scene division.
[0071] Step S204307, using the logical relationship words as second segmentation points, segmenting the text content to obtain a plurality of text segments;
[0072] Specifically, the text content is divided based on these logical relationship words to generate a series of text paragraphs around specific logical relationships. Each text paragraph represents a logical unit in the scene, which can be an explanation of a causal relationship, a fragment of a progressive process, or an explanation of a conditional structure.
[0073] Step S204308: determining the target virtual scene corresponding to each of the text segments as the plurality of virtual sub-scenes corresponding to the target virtual scene.
[0074] Specifically, the virtual scene parts corresponding to each text paragraph are clearly defined as independent virtual sub-scenes. This process is essentially to modularize the target virtual scene according to its internal logical structure. Each text paragraph, whether it involves a detailed analysis of cause and effect, the gradual advancement of progressive events, or the multi-dimensional consideration of conditional structures, is regarded as a unit that carries specific logical content and is then transformed into an independent virtual sub-scene.
[0075] In order to obtain the virtual sub-scene, in an optional implementation manner, the division method of the target virtual scene is determined according to the type tag, and the target virtual scene is divided according to the division method to obtain a plurality of the virtual sub-scenes, and the step S2043 further includes:
[0076] Step S204309, when the type tag corresponding to the target virtual scene is the multi-element type tag, all target elements in the target virtual scene are obtained, where the target elements include people, events, places, objects and concepts;
[0077] Specifically, extract the target virtual scene and extract all target elements contained in the scene. These elements cover multi-dimensional contents such as characters, events, locations, items, and concepts. Character elements may be historical figures, witnesses of important events, or characters related to the target virtual scene; event elements may be major historical events, events in the field of physics, etc. related to the target virtual scene; location elements may be important meeting places, landmarks, etc. indicating locations; item elements may involve items such as souvenirs and materials; concept elements may involve conceptual explanations of a specific person or event, etc. Through this refined element classification and extraction, the system can build an element library that comprehensively reflects the characteristics of the virtual scene of multiple element types, providing a basis for subsequent scene division.
[0078] Step S204310, determine the importance of the above target elements in the above target virtual scene, and obtain the importance weights corresponding to each of the above target elements;
[0079] Specifically, conduct importance assessments on various target elements (including characters, events, locations, items, and concepts) extracted from the scene, and accordingly assign a quantified importance weight to each element, judge the role of each element in the scene, and then determine its importance.
[0080] Step S204311, based on a preset value, combine the above target elements to obtain target element groups. The sum of the above importance weights corresponding to each of the above target elements in the above target element groups is less than the above preset value, and at least one of the above characters, one of the above events, one of the above locations, one of the above items, and one of the above concepts are included in the above target element groups;
[0081] Specifically, combine various target elements (such as characters, events, locations, items, and concepts) extracted to create target element groups. This process ensures that the elements within each combination have appropriate importance weights, and the sum of the weights does not exceed the preset threshold, thereby balancing the richness of the scene content and the concentration of information. When constructing each target element group, at least one element from each category (characters, events, locations, items, and concepts) is included to construct a comprehensive and multi-dimensional virtual experience.
[0082] Step S204312, determine the above target virtual scenes corresponding to each of the above target element groups as multiple above virtual sub-scenes corresponding to the above target virtual scene.
[0083] Specifically, regard each target element group as an independent narrative unit. Each unit includes multiple elements such as characters, events, locations, items, and concepts. In this way, the original target virtual scene is transformed into a series of virtual sub-scenes focused on specific themes and element combinations.
[0084] In an alternative embodiment, to provide a better VR experience for the target user, after obtaining the final virtual scene, the above method further includes:
[0085] Step S301, when the target user uses a VR device to view the final virtual scene, control the VR system to stop playing at the time point where the interaction node is located, so that the target user can use the interaction node to interact;
[0086] Specifically, when the experience time point of the target user coincides with the time coordinate where the interaction node is located, the system will automatically pause the playback process of the VR content, providing the user with an interaction node to deeply interact with the specific interaction node through multi-modal interaction methods such as a handle, gestures, or voice. For example, the user can click on an icon of a historical event to deeply understand the event background and impact;
[0087] Step S302, when the target user does not use the interaction node to interact within a predetermined time, control the VR system to automatically play the final virtual scene.
[0088] Specifically, when it is detected that the target user does not perform any interaction behavior within a preset waiting time (for example, within 30 seconds to 1 minute) after reaching the interaction node, the system will automatically evaluate the current attention state of the user. If it is determined that the user may not interact temporarily due to high immersion, in-depth thinking, or unfamiliar operation, etc., the system will automatically resume the playback of the scene to avoid the interruption of the learning experience or the decline of learning efficiency caused by a long pause state.
[0089] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0090] The embodiment of the present application also provides a VR teaching device. It should be noted that a VR teaching device in the embodiment of the present application can be used to execute the VR teaching method provided in the embodiment of the present application. The device for implementing the above embodiment and the preferred embodiment has been described and will not be repeated here. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0091] The following introduces a VR teaching device provided by the embodiment of the present application.
[0092] Figure 3It is a structural block diagram of a VR teaching device according to an embodiment of the present application. As Figure 3 shown, the device includes:
[0093] An acquisition unit 10, configured to acquire target instruction information and preference information of a target user, where the target instruction information is query information sent by the target user to the VR system, and the preference information includes visual preference, auditory preference, and interaction mode preference;
[0094] Specifically, first, receive the query information input by the target user through the VR interface or an external device, that is, the target instruction information, which covers the knowledge topics and content details that the user expects to explore or master in depth. At the same time, receive the preference information of the target user. When the user first wears the VR device, the system collects the preference information through the VR device, including visual preferences, such as preferences for realistic style, cartoon style, or abstract expression, as well as personal tendencies for color and brightness; auditory preferences, involving the choice of background music, the intonation and rhythm preferences of voice explanations; and interaction mode preferences. The user may be more inclined to use a handle, voice commands, or gesture control to communicate with the virtual environment. When the user uses VR multiple times, the system automatically collects the user's historical data, including the user's browsing history, search records, preference information, and data in historical virtual reality scenarios in the VR system. When the user uses the VR device again, through the unique identifier of the VR device, such as ID, ID card information, mobile phone number, etc., automatically obtain the user's personal preference information, historical browsing records, search records, and other information.
[0095] A construction unit 20, configured to construct a scene content generation model, where the scene content generation model is used to generate a virtual scene of the VR system according to the target instruction information and the preference information, where the scene content generation model is trained using multiple sets of training data, and each set of training data in the multiple sets of training data includes historical target instruction information and historical preference information;
[0096] Specifically, by constructing a scene content generation model, the core function of this model is to intelligently analyze and fuse the target instruction information and preference information of the target user, so as to generate a highly customized VR virtual scene. The training process of this model uses a large number of historical data sets. Each data set contains the query information sent by past users (historical target instruction information), as well as the personal preferences of these users in terms of vision, hearing, and interaction mode (historical preference information). Through deep learning technology, the model can extract the correlation patterns between user instructions and preferences, and how these patterns affect the content design and presentation of virtual scenes. Finally, the scene content generation model can generate personalized scenes that meet the requirements according to the target instruction information and preference information of the user.
[0097] A training unit 30 for inputting the above-mentioned target instruction information and the above-mentioned preference information into the above-mentioned scenario content generation model to obtain a target virtual scenario, so that the target virtual scenario meets the requirements of the above-mentioned preference information of the above-mentioned target user;
[0098] Specifically, by combining the query information proposed by the target user, that is, the target instruction information, with the user's personalized preference information, a virtual learning scenario that meets the specific needs of the user is intelligently generated. Specifically, the collected target instruction information and preference information are used as input parameters and transmitted to the trained scenario content generation model. The model uses the patterns it has learned to analyze the user's intention, and at the same time considers the user's preferences in terms of vision, hearing, and interaction methods, and accurately generates a target virtual scenario that highly matches the user's needs. Whether the user pursues a visually rich and vivid visual effect, or prefers an immersive auditory experience, or has a preference for a specific interaction method, the model can generate a virtual scenario that meets these preferences, ensuring that each user can obtain the best learning experience.
[0099] A partitioning unit 40 for partitioning the above-mentioned target virtual scenario according to the type of the above-mentioned target virtual scenario to obtain a plurality of virtual sub-scenarios, and adding interaction nodes to each of the above-mentioned virtual sub-scenarios, where the interaction nodes are used to answer the above-mentioned target instruction information related to the corresponding above-mentioned virtual sub-scenario or jump to the above-mentioned virtual sub-scenario corresponding to the above-mentioned target instruction information that is not related to the corresponding above-mentioned virtual sub-scenario;
[0100] Specifically, according to the type of the generated target virtual scenario, it is subdivided into a plurality of smaller and more manageable virtual sub-scenarios. This partitioning process aims to identify the key themes or teaching units in the scenario, ensuring that each sub-scenario can focus on specific learning objectives and provide in-depth teaching content. More importantly, after obtaining the virtual sub-scenarios by partitioning, interaction nodes will be embedded in each virtual sub-scenario. These node designs can respond to the learning needs of users within the sub-scenario. They can not only answer the target instruction information directly related to the current sub-scenario, guiding users to deeply understand the scenario content, but also provide a jump function, allowing users to quickly access other virtual sub-scenarios that are not directly related to the theme of the current virtual sub-scenario but are consistent with the user's overall learning objectives, realizing cross-domain exploration of knowledge.
[0101] A control unit 50 for using the above-mentioned target virtual scenario after adding the above-mentioned interaction nodes as the final virtual scenario and playing the above-mentioned final virtual scenario to teach the above-mentioned target user.
[0102] Specifically, after adding interaction nodes to the generated target virtual scene, it is transformed into the final virtual teaching scene, and then presented to the target user to guide them to carry out immersive learning. This process is not just simply playing the scene, but through the designed interaction nodes, closely integrating the user's learning path with the dynamic display of the virtual scene, creating an interactive learning experience.
[0103] Through this embodiment, in the above-mentioned VR teaching device, an acquisition unit is used to acquire the target instruction information and preference information of the target user. The above-mentioned target instruction information is the query information sent by the above-mentioned target user to the VR system, and the above-mentioned preference information includes visual preference, auditory preference, and interaction mode preference; a construction unit is used to construct a scene content generation model, and the above-mentioned scene content generation model is used to generate the virtual scene of the above-mentioned VR system according to the above-mentioned target instruction information and the above-mentioned preference information. Among them, the above-mentioned scene content generation model is trained using multiple groups of training data, and each group of the above-mentioned multiple groups of training data includes historical target instruction information and historical preference information; a training unit is used to input the above-mentioned target instruction information and the above-mentioned preference information into the above-mentioned scene content generation model to obtain a target virtual scene, so that the above-mentioned target virtual scene meets the requirements of the above-mentioned preference information of the above-mentioned target user; a division unit is used to divide the above-mentioned target virtual scene according to the type of the above-mentioned target virtual scene to obtain multiple virtual sub-scenes, and add interaction nodes to each of the above-mentioned virtual sub-scenes. The above-mentioned interaction nodes are used to answer the above-mentioned target instruction information related to the corresponding above-mentioned virtual sub-scene or jump to the above-mentioned virtual sub-scene corresponding to the above-mentioned target instruction information that is not related to the corresponding above-mentioned virtual sub-scene; a control unit is used to use the above-mentioned target virtual scene after adding the above-mentioned interaction nodes as the final virtual scene and play the above-mentioned final virtual scene to teach the above-mentioned target user. This application generates a target virtual scene that meets the needs of the target user according to the target instruction information and preference information of the target user, divides the target virtual scene to obtain multiple virtual sub-scenes, and sets interaction nodes for each virtual sub-scene, so that the user can obtain the target instruction information of the corresponding virtual sub-scene or jump to the virtual sub-scene corresponding to the target instruction information that is not related to the corresponding virtual sub-scene through the interaction nodes. The target virtual scene after adding the interaction nodes is the final virtual scene. By generating a personalized target virtual scene and interaction nodes according to the target instruction information and preference information of the target user, the problem that the content and interaction mode of the virtual scene are both fixed modes is avoided, the user can obtain a better VR experience, and the problem that the VR system in the prior art cannot meet the personalized needs of users is solved.
[0104] In an alternative embodiment, in order to obtain the scene content generation model, a scene content generation model is constructed. The above-mentioned construction unit includes:
[0105] The first construction module is used to obtain the multiple above-mentioned historical target instruction information and the multiple above-mentioned historical preference information;
[0106] The second construction module is used to determine the audience labels of the above-mentioned historical preference information, and the above-mentioned audience labels include education practitioner labels and ordinary user labels;
[0107] Specifically, collect the query information sent by multiple past users, that is, historical target instruction information, and these instruction information cover the learning needs of users for specific knowledge fields; at the same time, also record multiple historical preference information, and these information reflect the personal tendencies of users in terms of vision, hearing and interaction methods. Then, we conduct in-depth mining on each group of historical preference information to determine the matching audience labels, including education practitioner labels and ordinary user labels. This step aims to identify the professional background and personal attributes of users, so as to more accurately understand their learning behaviors and preference patterns.
[0108] The third construction module is used to extract feature vectors from each of the above-mentioned historical target instruction information and the above-mentioned historical preference information with the above-mentioned audience labels to obtain multiple historical target instruction vectors and multiple historical preference vectors;
[0109] The fourth construction module is used to train the LSTM neural network model with each of the above-mentioned historical target instruction vectors and each of the above-mentioned historical preference vectors to obtain a scene content generation model.
[0110] Specifically, a large number of collected historical target instruction information is subjected to feature vector extraction to extract feature vectors closely related to the user's query intention and scenario theme, namely historical target instruction vectors. At the same time, for historical preference information with clear audience labels (such as education practitioners, ordinary users, etc.), feature vector extraction is also performed to obtain historical preference vectors, which detail the specific preferences of users in terms of vision, hearing, and interaction methods. The specific extraction process for historical target instruction information is to clean the text in the historical target instruction information, remove stop words and special symbols, and then use word embedding technology to convert the processed text into a vector representation. Using a pre-trained word vector model, each word is mapped to a vector of a fixed dimension (such as 300 dimensions). The specific extraction process for historical preference information is to numerically encode each preference data. For example, the visual picture style can be divided into realistic (encoded as 0), cartoon (encoded as 1), abstract (encoded as 2), etc.; the color preference can be numerically encoded between 0 and 1 according to the hue range. For example, if the user prefers warm colors, it is encoded as a higher value, and if the user prefers cold colors, it is encoded as a lower value, etc. These encoded preference data are combined into a feature vector as another part of the input layer of the LSTM network. For example, after encoding the preference data, a 20-dimensional vector is formed. Subsequently, the system uses the extracted historical target instruction vectors and historical preference vectors to deeply train the LSTM (Long Short-Term Memory) neural network model. Through this training process, the model not only learns how to generate corresponding scenario content according to the user's instructions but also masters how to adjust the presentation method of the scenario according to the user's preferences, ensuring that the finally generated VR teaching scenario accurately reflects the user's needs and is highly customized to meet the user's personalized preferences.
[0111] In an alternative embodiment, in order to divide the target virtual sub-scene, the target virtual scene is divided according to the type of the target virtual scene described above to obtain a plurality of virtual sub-scenes. The dividing unit includes:
[0112] The first dividing module is used to construct a scene category model. Among them, the scene category model is trained using multiple sets of training data. Each set of training data in the multiple sets of training data includes historical target virtual scenes with type labels obtained within a historical time period. The type labels include time type labels, logical type labels, and multi-element type labels;
[0113] Specifically, the construction of the scenario category model is based on a large amount of historical target virtual scenario data. Each set of training data is labeled with scenario type tags, including time type tags, logical type tags, and multi-element type tags. These tags reflect the characteristics of the scenario in terms of time series, logical structure, and element combination. By using multiple sets of such labeled historical data to deeply train the model, the system can learn the inherent features and generation patterns of different types of scenarios, and thus can determine the type to which the target virtual scenario belongs.
[0114] A second partitioning module, configured to input the above-mentioned target virtual scenario into the above-mentioned scenario category model to obtain the above-mentioned type tag corresponding to the above-mentioned target virtual scenario;
[0115] Specifically, the initially generated target virtual scenario is used as input and passed to the deeply trained scenario category model. Based on its learning of a large number of historical target virtual scenarios, this model can quickly and accurately identify the core attributes of the target scenario and assign corresponding type tags, such as time type, logical type, or multi-element type.
[0116] A third partitioning module, configured to determine the partitioning method of the above-mentioned target virtual scenario according to the above-mentioned type tag, and partition the above-mentioned target virtual scenario according to the above-mentioned partitioning method to obtain multiple above-mentioned virtual sub-scenarios.
[0117] Specifically, the type tag output by the scenario category model is parsed, and this tag reveals the structural attributes of the target virtual scenario. Subsequently, based on these type tags, the most appropriate scenario partitioning strategy is automatically selected to cut the complex target scenario into a series of virtual sub-scenarios that are easier to understand and interact with. This process takes into account the structural partitioning of the scenario to ensure a high correlation for each virtual sub-scenario.
[0118] In order to obtain virtual sub-scenarios, in an alternative embodiment, the partitioning method of the above-mentioned target virtual scenario is determined according to the above-mentioned type tag, and the above-mentioned target virtual scenario is partitioned according to the above-mentioned partitioning method to obtain multiple above-mentioned virtual sub-scenarios. The above-mentioned third partitioning module includes:
[0119] A first sub-partitioning module, configured to, when the above-mentioned type tag corresponding to the above-mentioned target virtual scenario is the above-mentioned time type tag, extract a first time point and a second time point of the above-mentioned target virtual scenario by using a time expression recognition algorithm. The above-mentioned first time point is all the time points in the above-mentioned target virtual scenario, and the above-mentioned second time point is an important event node of the above-mentioned target virtual scenario;
[0120] Specifically, the time expression recognition algorithm is used to carefully sort out the time clues in the scene. In this process, the target virtual scene is first scanned comprehensively to accurately locate each time point, forming a complete first time point sequence, which covers all historical moments from the beginning to the end of the scene. Subsequently, the algorithm further conducts in-depth analysis to screen out the second time points with great significance or teaching value from the first time point sequence, that is, important event nodes, which are often associated with key teaching contents.
[0121] The second sub-module for partitioning is used to arrange the above-mentioned first time points and the above-mentioned second time points in chronological order to obtain a timeline.
[0122] Specifically, based on the first time points and the second time points, following the chronological order, a timeline is constructed to connect all historical moments and important events according to the time context, that is, the timeline carries the time points including important events.
[0123] The third sub-module for partitioning is used to divide the above-mentioned timeline with the above-mentioned second time points as the first dividing points to obtain multiple time periods.
[0124] Specifically, based on the second time points, that is, those time nodes closely related to important events or teaching key points, as the first dividing points. Based on these first dividing points, the complete timeline is accurately divided, thus generating a series of time periods with clear key stages or teaching themes.
[0125] The fourth sub-module for partitioning is used to determine the above-mentioned target virtual scenes corresponding to each time period as multiple above-mentioned virtual sub-scenes corresponding to the above-mentioned target virtual scene.
[0126] Specifically, each time period obtained by subdividing the timeline is regarded as an independent segment of the target virtual scene, and the virtual sub-scene corresponding to each time period is determined as an independent virtual sub-scene.
[0127] In order to obtain virtual sub-scenes, in an optional implementation manner, the partitioning method of the above-mentioned target virtual scene is determined according to the above-mentioned type label, and the above-mentioned target virtual scene is partitioned according to the above-mentioned partitioning method to obtain multiple above-mentioned virtual sub-scenes. The above-mentioned third partitioning module further includes:
[0128] The fifth sub-module for partitioning is used to obtain all the text contents of the above-mentioned target virtual scene when the above-mentioned type label corresponding to the above-mentioned target virtual scene is the above-mentioned logical type label.
[0129] Specifically, the target virtual scene is deeply scanned to ensure that all text content in the scene is captured. This process not only covers the statically displayed text information, but also integrates the text data that may be generated in dynamic interaction according to the actual scene to form a text content collection, providing a data basis for subsequent scene segmentation.
[0130] A sixth division submodule is used to identify the logical relationship words of the above text content, wherein the above logical relationship words include causal relationship words, progressive relationship words and conditional relationship words;
[0131] Specifically, natural language processing technology is used to extract logical relationship words in text content. These logical relationship words include but are not limited to causal relationship words (such as "therefore", "due to", "result"), progressive relationship words (such as "then", "afterwards", "further") and conditional relationship words (such as "if", "unless", "however"). They are like hidden threads of story narration, revealing the logical connection and advancement context between the various parts of the scene content. By locating these logical relationship words, it is possible to understand the narrative logic of the scene, distinguish the causal relationship between different elements, the continuous steps of event development, and the multiple paths triggered by conditions, thereby providing guidance information for subsequent scene division.
[0132] A seventh segmentation submodule is used to segment the text content using the logical relationship words as the second segmentation point to obtain a plurality of text segments;
[0133] Specifically, the text content is divided based on these logical relationship words to generate a series of text paragraphs around specific logical relationships. Each text paragraph represents a logical unit in the scene, which can be an explanation of a causal relationship, a fragment of a progressive process, or an explanation of a conditional structure.
[0134] The eighth division submodule is used to determine the target virtual scene corresponding to each of the text segments into a plurality of virtual sub-scenes corresponding to the target virtual scene.
[0135] Specifically, the virtual scene parts corresponding to each text paragraph are clearly defined as independent virtual sub-scenes. This process is essentially to modularize the target virtual scene according to its internal logical structure. Each text paragraph, whether it involves a detailed analysis of cause and effect, the gradual advancement of progressive events, or the multi-dimensional consideration of conditional structures, is regarded as a unit that carries specific logical content and is then transformed into an independent virtual sub-scene.
[0136] In order to obtain the virtual sub-scene, in an optional implementation manner, the division method of the target virtual scene is determined according to the type label, and the target virtual scene is divided according to the division method to obtain a plurality of the virtual sub-scenes, and the third division module further includes:
[0137] The ninth partitioning sub-module is used to obtain all target elements in the target virtual scene when the type label corresponding to the target virtual scene is the multi-element type label, and the target elements include people, events, locations, items, and concepts;
[0138] Specifically, extract the target virtual scene to extract all target elements contained in the scene, and these elements cover multi-dimensional contents such as people, events, locations, items, and concepts. The person element may be a historical figure, a witness to an important event, or a person related to the target virtual scene; the event element may be a major historical event, an event in the field of physics, etc., related to the target virtual scene; the location element may be an important meeting location, a landmark, etc., indicating the location element; the item element may involve items such as souvenirs and materials; the concept element may involve the concept explanation of a specific person or event, etc. Through this fine element classification and extraction, the system can build an element library that comprehensively reflects the virtual scene characteristics of the multi-element type, providing a basis for subsequent scene partitioning.
[0139] The tenth partitioning sub-module is used to determine the importance level of the above target elements in the above target virtual scene and obtain the importance weight corresponding to each above target element;
[0140] Specifically, evaluate the importance of various target elements (including people, events, locations, items, and concepts) extracted from the scene, and accordingly assign a quantified importance weight to each element, judge the role of each element in the scene, and then determine its importance level.
[0141] The eleventh partitioning sub-module is used to combine the above target elements based on a preset value to obtain a target element group, and the sum of the above importance weights corresponding to each above target element in the above target element group is less than the above preset value, and at least one of the above people, one of the above events, one of the above locations, one of the above items, and one of the above concepts are included in the above target element group;
[0142] Specifically, combine various target elements (such as people, events, locations, items, and concepts) extracted to create a target element group. This process ensures that the elements within each combination have appropriate importance weights, and the sum of the weights does not exceed the preset threshold, thereby balancing the richness of the scene content and the concentration of information. When constructing each target element group, at least one element from each category (people, events, locations, items, and concepts) is included to construct a comprehensive and multi-dimensional virtual experience.
[0143] The twelfth partitioning sub-module is used to determine the above-mentioned target virtual scenes corresponding to each of the above-mentioned target element groups as a plurality of the above-mentioned virtual sub-scenes corresponding to the above-mentioned target virtual scene.
[0144] Specifically, each target element group is regarded as an independent narrative unit, and each unit includes multiple elements such as characters, events, locations, items, and concepts. In this way, the original target virtual scene is transformed into a series of virtual sub-scenes focusing on specific themes and element combinations.
[0145] To enable a better VR experience for the target user, in an optional implementation manner, after obtaining the final virtual scene, the above method further includes:
[0146] The first interaction unit is used to control the VR system to stop playing at the time point where the interaction node is located when the above-mentioned target user views the above-mentioned final virtual scene using a VR device, so that the above-mentioned target user can use the above-mentioned interaction node for interaction;
[0147] Specifically, when the experience time point of the target user coincides with the time coordinate where the interaction node is located, the system will automatically pause the playback process of the VR content, providing the user with an interaction node to deeply interact with a specific interaction node through multi-modal interaction methods such as a handle, gestures, or voice. For example, the user can click on an icon of a historical event to deeply understand the event background and impact;
[0148] The second interaction unit is used to control the VR system to automatically play the above-mentioned final virtual scene when the above-mentioned target user does not use the above-mentioned interaction node for interaction within a predetermined time.
[0149] Specifically, when it is detected that the target user has not performed any interaction behavior within a preset waiting time (for example, within 30 seconds to 1 minute) after reaching the interaction node, the system will automatically evaluate the current attention state of the user. If it is determined that the user may not interact temporarily due to high immersion, in-depth thinking, or unfamiliar operation, etc., the system will automatically resume the playback of the scene to avoid the interruption of the learning experience or the decline of learning efficiency caused by a long pause state.
[0150] The above-mentioned VR teaching device includes a processor and a memory. The above-mentioned acquisition unit, construction unit, training unit, partitioning unit, and control unit, etc. are all stored in the memory as program units, and the processor executes the above-mentioned program units stored in the memory to implement corresponding functions. The above-mentioned modules are all located in the same processor; or, the above-mentioned each module is located in different processors in any combination form.
[0151] The processor contains a kernel, which retrieves corresponding program units from the memory. One or more kernels can be set, and the generation quality of the final virtual scene can be improved by adjusting the kernel parameters.
[0152] The memory may include non-permanent memory in a computer-readable medium, forms such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one memory chip.
[0153] An embodiment of the present invention provides a computer-readable storage medium, and the above computer-readable storage medium includes a stored program. Among them, when the above program runs, it controls the device where the above computer-readable storage medium is located to execute the above VR teaching method.
[0154] Specifically, a VR teaching method includes:
[0155] Step S201, obtaining target instruction information and preference information of a target user. The above target instruction information is query information sent by the above target user to the VR system, and the above preference information includes visual preference, auditory preference, and interaction mode preference;
[0156] Step S202, constructing a scene content generation model. The above scene content generation model is used to generate the virtual scene of the above VR system according to the above target instruction information and the above preference information. Among them, the above scene content generation model is trained using multiple groups of training data, and each group of training data in the above multiple groups of training data includes historical target instruction information and historical preference information;
[0157] Step S203, inputting the above target instruction information and the above preference information into the above scene content generation model to obtain a target virtual scene, so that the above target virtual scene meets the requirements of the above preference information of the above target user;
[0158] Step S204, dividing the above target virtual scene according to the type of the above target virtual scene to obtain multiple virtual sub-scenes, and adding interaction nodes to each of the above virtual sub-scenes. The above interaction nodes are used to answer the above target instruction information related to the corresponding above virtual sub-scene or jump to the above virtual sub-scene corresponding to the above target instruction information that is not related to the corresponding above virtual sub-scene;
[0159] Step S205, using the above target virtual scene after adding the above interaction nodes as the final virtual scene, and playing the above final virtual scene to teach the above target user.
[0160] An embodiment of the present invention provides a processor, which is used to run a program. When the program runs, it executes the above VR teaching method.
[0161] Specifically, a VR teaching method includes:
[0162] Step S201: Obtain the target instruction information and preference information of the target user. The target instruction information is the query information sent by the target user to the VR system. The preference information includes visual preference, auditory preference, and interaction mode preference;
[0163] Step S202: Construct a scene content generation model, which is used to generate the virtual scene of the VR system according to the target instruction information and the preference information. The scene content generation model is trained using multiple sets of training data, and each set of training data in the multiple sets of training data includes historical target instruction information and historical preference information;
[0164] Step S203: Input the target instruction information and the preference information into the scene content generation model to obtain a target virtual scene, so that the target virtual scene meets the requirements of the preference information of the target user;
[0165] Step S204: Divide the target virtual scene according to the type of the target virtual scene to obtain multiple virtual sub-scenes, and add interaction nodes to each virtual sub-scene. The interaction nodes are used to answer the target instruction information related to the corresponding virtual sub-scene or jump to the virtual sub-scene corresponding to the target instruction information not related to the corresponding virtual sub-scene;
[0166] Step S205: Use the target virtual scene after adding the interaction nodes as the final virtual scene, and play the final virtual scene to teach the target user.
[0167] An embodiment of the present application also provides a VR teaching system, including: one or more processors, a memory, and one or more programs. The one or more programs are stored in the memory and are configured to be executed by the one or more processors, including executing any one of the above VR teaching methods.
[0168] Specifically, a VR teaching method includes:
[0169] Step S201: Obtain the target instruction information and preference information of the target user. The target instruction information is the query information sent by the target user to the VR system. The preference information includes visual preference, auditory preference, and interaction mode preference;
[0170] Step S202: Build a scenario content generation model. The above scenario content generation model is used to generate the virtual scenario of the above VR system according to the above target instruction information and the above preference information. Among them, the above scenario content generation model is trained using multiple sets of training data, and each set of training data in the above multiple sets of training data includes historical target instruction information and historical preference information;
[0171] Step S203: Input the above target instruction information and the above preference information into the above scenario content generation model to obtain a target virtual scenario, so that the above target virtual scenario meets the requirements of the above preference information of the above target user;
[0172] Step S204: Divide the above target virtual scenario according to the type of the above target virtual scenario to obtain multiple virtual sub-scenarios, and add interaction nodes to each of the above virtual sub-scenarios. The above interaction nodes are used to answer the above target instruction information related to the corresponding above virtual sub-scenario or jump to the above virtual sub-scenario corresponding to the above target instruction information that is not related to the corresponding above virtual sub-scenario;
[0173] Step S205: Use the above target virtual scenario after adding the above interaction nodes as the final virtual scenario, and play the above final virtual scenario to teach the above target user.
[0174] Obviously, those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program code executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the present invention is not limited to any specific combination of hardware and software.
[0175] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0176] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks, or means for implementing the functions specified in one block or multiple blocks.
[0177] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks, or means for implementing the functions specified in one block or multiple blocks.
[0178] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks, or means for implementing the functions specified in one block or multiple blocks.
[0179] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.
[0180] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory. The memory is an example of computer-readable media.
[0181] A computer-readable medium includes permanent and non-permanent, removable and non-removable media that can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, magnetic tape disk storage, or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0182] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.
[0183] From the above description, it can be seen that the above embodiments of the present application achieve the following technical effects:
[0184] 1), A VR teaching method of the present application generates a target virtual scene that meets the needs of the target user according to the target instruction information and preference information of the target user, divides the target virtual scene to obtain a plurality of virtual sub-scenes, and sets interaction nodes for each virtual sub-scene, so that the user can obtain the target instruction information of the corresponding virtual sub-scene or jump to the virtual sub-scene corresponding to the target instruction information that is not related to the corresponding virtual sub-scene through the interaction nodes. The target virtual scene after adding the interaction nodes is the final virtual scene. By generating a personalized target virtual scene and interaction nodes according to the target instruction information and preference information of the target user, the problem that the content and interaction method of the virtual scene are both in a fixed mode is avoided, and the user can obtain a better VR experience, solving the problem that the VR system in the prior art cannot meet the personalized needs of users.
[0185] 2) A VR teaching device of the present application generates a target virtual scene that meets the needs of the target user according to the target instruction information and preference information of the target user, divides the target virtual scene to obtain multiple virtual sub-scenes, and sets interaction nodes for each virtual sub-scene, so that the user can obtain the target instruction information of the corresponding virtual sub-scene or jump to the virtual sub-scene corresponding to the target instruction information that is not related to the corresponding virtual sub-scene through the interaction nodes. The target virtual scene after adding the interaction nodes is the final virtual scene. By generating a personalized target virtual scene and interaction nodes according to the target instruction information and preference information of the target user, the problem that the content and interaction method of the virtual scene are both in a fixed mode is avoided, the user can obtain a better VR experience, and the problem that the VR system in the prior art cannot meet the personalized needs of users is solved.
[0186] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A VR teaching method, characterized in that: include: Acquire target instruction information and preference information of a target user, wherein the target instruction information is query information sent by the target user to the VR system, and the preference information includes visual preference, auditory preference, and interaction mode preference; Constructing a scene content generation model, wherein the scene content generation model is used to generate a virtual scene of the VR system according to the target instruction information and the preference information, wherein the scene content generation model is trained using multiple sets of training data, and each set of training data in the multiple sets of training data includes historical target instruction information and historical preference information; Inputting the target instruction information and the preference information into the scene content generation model to obtain a target virtual scene, so that the target virtual scene meets the requirement of the preference information of the target user; Divide the target virtual scene according to the type of the target virtual scene to obtain a plurality of virtual sub-scenes, and add an interactive node for each of the virtual sub-scenes, wherein the interactive node is used to answer the target instruction information related to the corresponding virtual sub-scene or jump to the virtual sub-scene corresponding to the target instruction information not related to the corresponding virtual sub-scene; The target virtual scene after adding the interactive node is used as the final virtual scene, and the final virtual scene is played to teach the target user.
2. The method according to claim 1, characterized in that Build a scene content generation model, including: Acquire a plurality of the historical target instruction information and a plurality of the historical preference information; Determine an audience tag for each of the historical preference information, wherein the audience tag includes an education practitioner tag and a general user tag; Extracting feature vectors from each of the historical target instruction information and each of the historical preference information having the audience tag to obtain a plurality of historical target instruction vectors and a plurality of historical preference vectors; The LSTM neural network model is trained using each of the historical target instruction vectors and each of the historical preference vectors to obtain a scene content generation model.
3. The method according to claim 1, characterized in that The target virtual scene is divided according to the type of the target virtual scene to obtain a plurality of virtual sub-scenes, including: Constructing a scene category model, wherein the scene category model is obtained by training using multiple sets of training data, each set of training data in the multiple sets of training data includes historical target virtual scenes with type labels obtained within a historical time period, and the type labels include time type labels, logic type labels, and multi-element type labels; Inputting the target virtual scene into the scene category model to obtain the type label corresponding to the target virtual scene; The division method of the target virtual scene is determined according to the type label, and the target virtual scene is divided according to the division method to obtain a plurality of virtual sub-scenes.
4. The method according to claim 3, characterized in that Determining a division method of the target virtual scene according to the type tag, and dividing the target virtual scene according to the division method to obtain a plurality of virtual sub-scenes, including: In the case where the type tag corresponding to the target virtual scene is the time type tag, a time expression recognition algorithm is used to extract a first time point and a second time point of the target virtual scene, wherein the first time point is all time points in the target virtual scene, and the second time point is an important event node of the target virtual scene; Arrange the first time point and the second time point in chronological order to obtain a time axis; Splitting the time axis with the second time point as a first split point to obtain multiple time periods; The target virtual scene corresponding to each time period is determined as a plurality of virtual sub-scenes corresponding to the target virtual scene.
5. The method according to claim 3, characterized in that: Determining a division method of the target virtual scene according to the type tag, and dividing the target virtual scene according to the division method to obtain a plurality of virtual sub-scenes, further comprising: When the type tag corresponding to the target virtual scene is the logical type tag, acquiring all text contents of the target virtual scene; Identify logical relationship words of the text content, wherein the logical relationship words include causal relationship words, progressive relationship words and conditional relationship words; Using the logical relationship words as second segmentation points, segmenting the text content to obtain a plurality of text segments; The target virtual scene corresponding to each of the text segments is determined as a plurality of virtual sub-scenes corresponding to the target virtual scene.
6. The method according to claim 3, characterized in that Determining a division method of the target virtual scene according to the type tag, and dividing the target virtual scene according to the division method to obtain a plurality of virtual sub-scenes, further comprising: In a case where the type tag corresponding to the target virtual scene is the multi-element type tag, acquiring all target elements in the target virtual scene, the target elements including characters, events, places, objects, and concepts; Determine the importance of the target element in the target virtual scene, and obtain the importance weight corresponding to each target element; Based on a preset value, the target elements are combined to obtain a target element group, the sum of the importance weights corresponding to the target elements in the target element group is less than the preset value, and the target element group includes at least one person, one event, one place, one object and one concept; The target virtual scene corresponding to each target element group is determined as a plurality of virtual sub-scenes corresponding to the target virtual scene.
7. The method according to claim 1, characterized in that After obtaining the final virtual scene, the method further includes: When the target user uses a VR device to view the final virtual scene, controlling the VR system to stop playing at the time point where the interactive node is located, so that the target user uses the interactive node to interact; When the target user does not use the interactive node to interact within a predetermined time, the VR system is controlled to automatically play the final virtual scene.
8. A VR teaching device, characterized in that: include: An acquisition unit, used to acquire target instruction information and preference information of a target user, wherein the target instruction information is query information sent by the target user to the VR system, and the preference information includes visual preference, auditory preference and interaction mode preference; A construction unit, configured to construct a scene content generation model, wherein the scene content generation model is configured to generate a virtual scene of the VR system according to the target instruction information and the preference information, wherein the scene content generation model is trained using a plurality of sets of training data, and each set of training data in the plurality of sets of training data includes historical target instruction information and historical preference information; A training unit, used for inputting the target instruction information and the preference information into the scene content generation model to obtain a target virtual scene, so that the target virtual scene meets the requirement of the preference information of the target user; A division unit, used for dividing the target virtual scene according to the type of the target virtual scene to obtain a plurality of virtual sub-scenes, and adding an interactive node to each of the virtual sub-scenes, wherein the interactive node is used to answer the target instruction information related to the corresponding virtual sub-scene or jump to the virtual sub-scene corresponding to the target instruction information not related to the corresponding virtual sub-scene; The control unit is used to use the target virtual scene after adding the interactive node as the final virtual scene, and play the final virtual scene to teach the target user.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 7.
10. A VR teaching system, characterized in that: include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include methods for executing any one of claims 1 to 7.