Roaming video generation method and device, equipment and storage medium

By generating roaming video call instructions that include the travel route, panoramic video data, and narration information, the problem of fixed node call order is solved, enabling flexible roaming video display.

CN121531203APending Publication Date: 2026-02-13PEKING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511471161.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-15
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

In existing methods of generating roaming videos, the order in which roaming video nodes are called is fixed, and the configuration file is strongly coupled with the calling program, making it cumbersome to change playback settings.

Method used

Generate invocation instructions for the roaming video. The instructions include the configuration file of the target node, which contains the travel route, panoramic video data, narrator image information, narration information, and 3D model. The instructions call and display this data to adjust the node display order.

Benefits of technology

It allows users to adjust the display order of nodes according to their needs. The displayed content can be changed simply by modifying the configuration file. The operation is very simple.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121531203A_ABST
    Figure CN121531203A_ABST
Patent Text Reader

Abstract

The invention provides a roaming video generation method and device, equipment and a storage medium, and the method comprises the steps: generating a calling instruction for a roaming video, and enabling the calling instruction to comprise a called target node; calling a target configuration file corresponding to the target node based on the calling instruction; playing an advancing route from the first preset node to the target node; and displaying the panoramic video data corresponding to the target node, the explanation human image information corresponding to the target node, the explanation information of the target node and the target three-dimensional model corresponding to the target node. According to the embodiment of the invention, the corresponding configuration files are configured for the plurality of nodes of the roaming video, the display sequence of the nodes can be adjusted based on user requirements, the displayed roaming video can be changed only by changing the configuration files corresponding to the roaming video, and the operation is simple.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, specifically relating to a method, apparatus, device, and storage medium for generating roaming videos. Background Technology

[0002] With the continuous development of various sensors, computing infrastructure and artificial intelligence, media technology is undergoing a leapfrog transformation from flat video to immersive video. Roaming systems are a killer application of immersive video technology, providing users with an immersive roaming experience in immersive video space.

[0003] However, in the current method of generating roaming videos, the order of calling roaming video nodes is relatively fixed, and the roaming video configuration file and the calling program are strongly coupled. When it is necessary to change the roaming video to be played, the roaming video configuration file and the calling program need to be reconfigured, which is cumbersome. Summary of the Invention

[0004] This application proposes a method, apparatus, device, and storage medium for generating roaming videos, which can solve the technical problems in current roaming video generation methods, such as the fixed order of calling roaming video nodes, the strong coupling between the roaming video configuration file and the calling program, and the need to reconfigure the roaming video configuration file and the calling program when changing the roaming video to be played, which is cumbersome.

[0005] The first aspect of this application proposes a method for generating roaming videos, including: Generate a calling instruction for the roaming video, the calling instruction including the target node to be called, the roaming video including multiple nodes, each node corresponding to a configuration file; The target configuration file corresponding to the target node is invoked based on the invocation instruction. The target configuration file includes the travel route from the first preset node to the target node, the panoramic video data corresponding to the target node, the image information of multiple narrators corresponding to the target node, the narration information of the target node, and the target 3D model corresponding to the target node. Play the travel route from the first preset node to the target node; The system displays the panoramic video data corresponding to the target node, the narrator's image information corresponding to the target node, the narration information of the target node, and the target 3D model corresponding to the target node.

[0006] An embodiment of the second aspect of this application provides a roaming video generation apparatus, comprising: The generation module is used to generate a calling instruction for the roaming video. The calling instruction includes the target node to be called. The roaming video includes multiple nodes, and each node corresponds to a configuration file. The calling module is used to call the target configuration file corresponding to the target node based on the calling instruction. The target configuration file includes the travel route from the first preset node to the target node, the panoramic video data corresponding to the target node, the image information of multiple narrators corresponding to the target node, the narration information of the target node, and the target 3D model corresponding to the target node. The playback module is used to play the travel route from the first preset node to the target node; The display module is used to display the panoramic video data corresponding to the target node, the image information of the narrator corresponding to the target node, the narration information of the target node, and the target 3D model corresponding to the target node.

[0007] An embodiment of the third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, the processor running the computer program to implement the method described in the first aspect above.

[0008] An embodiment of the fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the method described in the first aspect above.

[0009] The technical solutions provided in this application embodiment have at least the following technical effects or advantages: This application proposes a method, apparatus, device, and storage medium for generating roaming videos, comprising: generating a calling instruction for the roaming video, the calling instruction including a target node to be called, the roaming video including multiple nodes, each node corresponding to a configuration file; calling the target configuration file corresponding to the target node based on the calling instruction, the target configuration file including a travel route from a first preset node to the target node, panoramic video data corresponding to the target node, image information of multiple narrators corresponding to the target node, narration information of the target node, and a target 3D model corresponding to the target node; playing the travel route from the first preset node to the target node; and displaying the panoramic video data corresponding to the target node, the image information of the narrators corresponding to the target node, the narration information of the target node, and the target 3D model corresponding to the target node. This application embodiment, by configuring corresponding configuration files for multiple nodes of the roaming video respectively, allows for adjustment of the display order of nodes based on user needs, and only requires changing the configuration file corresponding to the roaming video to change the displayed roaming video, making the operation simple.

[0010] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0011] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 A flowchart of a roaming video generation method according to an embodiment of this application is shown; Figure 2 A flowchart of a roaming video generation method according to an embodiment of this application is shown; Figure 3 This illustration shows a schematic diagram of the structure of a roaming video generation apparatus according to an embodiment of this application; Figure 4 This illustration shows a schematic diagram of the structure of an electronic device according to an embodiment of this application; Figure 5 A schematic diagram of a storage medium provided in one embodiment of this application is shown. Detailed Implementation

[0012] Exemplary embodiments of this application will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of this application are shown in the drawings, it should be understood that this application may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of this application and to fully convey the scope of this application to those skilled in the art.

[0013] It should be noted that, unless otherwise stated, the technical or scientific terms used in this application shall have the ordinary meaning as understood by one of ordinary skill in the art to which this application pertains.

[0014] The roaming video generation method of this application can be executed by a computing device. The computing device can be a server, such as a single server, multiple servers, a server cluster, a cloud computing platform, etc. Optionally, the computing device can also be a terminal device, such as a mobile phone, tablet computer, game console, portable computer, desktop computer, advertising machine, all-in-one machine, etc. This application does not limit the type or number of computing devices.

[0015] Based on the aforementioned background technology, in the current method of generating roaming videos, the roaming video configuration file and the calling program are highly coupled. That is, various contents in the roaming video configuration file are written in the calling program, which results in a relatively fixed calling order of roaming video nodes. When changes are needed, the calling program needs to be modified, which leads to cumbersome operation.

[0016] To address the aforementioned issues, this application proposes a method, apparatus, device, and storage medium for generating roaming videos. The method includes: generating a calling instruction for the roaming video, the calling instruction including a target node to be called; the roaming video comprising multiple nodes, each node corresponding to a configuration file; calling the target configuration file corresponding to the target node based on the calling instruction; the target configuration file including a travel route from a first preset node to the target node, panoramic video data corresponding to the target node, image information of multiple narrators corresponding to the target node, narration information of the target node, and a target 3D model corresponding to the target node; playing the travel route from the first preset node to the target node; and displaying the panoramic video data corresponding to the target node, the image information of the narrators corresponding to the target node, the narration information of the target node, and the target 3D model corresponding to the target node. This application's embodiment, by configuring corresponding configuration files for each node of the roaming video, allows for adjustment of the node display order based on user needs. Furthermore, only the configuration file corresponding to the roaming video needs to be changed to modify the displayed roaming video, making the operation simple.

[0017] The following describes a method for generating roaming videos according to an embodiment of this application, with reference to the accompanying drawings.

[0018] See Figure 1 The method specifically includes the following steps: S101. Generate a call command for the roaming video.

[0019] The invocation command includes the target node to be invoked. The roaming video includes multiple nodes, each corresponding to a configuration file.

[0020] Among them, roaming videos can be roaming videos of landscape scenes, roaming videos of indoor scenes, and roaming videos of game scenes, etc.

[0021] Each node can correspond to a scene node in the roaming video. For example, in a landscape scene roaming video, each node can correspond to a landscape, and in an indoor scene roaming video, each node can correspond to a room, and so on.

[0022] Each node can correspond to a configuration file, which may include the travel route from the first preset node to the node, the panoramic video data corresponding to the node, the narrator's image information corresponding to the node, the narration information of the node, and the target 3D model corresponding to the node.

[0023] In some embodiments, each roaming video may correspond to a configuration file, and each node may correspond to a sub-configuration file within that configuration file.

[0024] In some embodiments, the configuration file of the node can be invoked based on the node identifier.

[0025] The invocation command can be generated based on the received user input information or based on a preset strategy.

[0026] For example, if the user input is "jump to landscape A" and the node identifier is identified as landscape A, then the corresponding calling command for landscape A can be generated.

[0027] The preset strategy can be to display the configuration file content corresponding to each node in a preset order. For example, if the current node is node A, after playing the configuration file content corresponding to node A, no instruction based on user input information is generated within a certain period of time, and in the preset order, node B is the next node after node A, then a call instruction including the identifier corresponding to node B is generated.

[0028] S102. Invoke the target configuration file corresponding to the target node based on the invocation command.

[0029] The target configuration file includes the travel route from the first preset node to the target node, the panoramic video data corresponding to the target node, the image information of multiple narrators corresponding to the target node, the narration information of the target node, and the target 3D model corresponding to the target node.

[0030] S103, Play the travel route from the first preset node to the target node.

[0031] S104. Display the panoramic video data corresponding to the target node, the image information of the narrator corresponding to the target node, the narration information of the target node, and the target 3D model corresponding to the target node.

[0032] After generating the invocation command, the invocation program can call the target configuration file corresponding to the target node and display the contents of the target configuration file.

[0033] First, the travel route from the first preset node to the target node can be played. The first preset node can be the node before the target node in the preset order, or it can be a fixed node selected from multiple nodes.

[0034] After the route playback is complete, the system can display the panoramic video data corresponding to the target node, the image information of the target narrator corresponding to the target node, the narration information of the target node, and the target 3D model corresponding to the target node.

[0035] The target narrator's image information can be any one of multiple narrator image information. The target narrator's image information can be a flat video or a dynamic mesh. The flat video format supports a transparent background. The narration information of the target node can include text information and the corresponding audio information. The target 3D model can be a 3D model corresponding to any one of the multiple objects corresponding to the target node, or a 3D model corresponding to a key object among the multiple objects corresponding to the target node.

[0036] This application proposes a method for generating a roaming video, comprising: generating a calling instruction for the roaming video, the calling instruction including a target node to be called, the roaming video including multiple nodes, each node corresponding to a configuration file; calling the target configuration file corresponding to the target node based on the calling instruction, the target configuration file including a travel route from a first preset node to the target node, panoramic video data corresponding to the target node, image information of multiple narrators corresponding to the target node, narration information of the target node, and a target 3D model corresponding to the target node; playing the travel route from the first preset node to the target node; and displaying the panoramic video data corresponding to the target node, the image information of the narrators corresponding to the target node, the narration information of the target node, and the target 3D model corresponding to the target node. This application embodiment, by configuring corresponding configuration files for multiple nodes of the roaming video respectively, allows for adjustment of the display order of nodes based on user needs, and only requires changing the configuration file corresponding to the roaming video to change the displayed roaming video, making the operation simple.

[0037] In some embodiments, generating a call instruction for roaming video includes: receiving a start instruction from the roaming system and, if no first input information from the user is received within a first preset time period, generating a call instruction for a second preset node corresponding to the start instruction based on a preset call order, wherein the first input information includes call information for a target node; or, receiving a call completion instruction from a first node and, if no first input information is received within a second preset time period, generating a call instruction for a third preset node corresponding to the first node based on a preset call order, wherein the first node is any of multiple nodes.

[0038] As mentioned above, the invocation command can be generated based on the received user input information or triggered based on a preset strategy.

[0039] In some embodiments, during the roaming system startup process, the process of generating the call instruction based on a preset strategy can be implemented as follows: receiving the roaming system startup instruction and not receiving the user's first input information within a first preset time period, generating the call instruction of the second preset node corresponding to the startup instruction based on a preset call order.

[0040] In general, after receiving the roaming system's start command, the roaming system is activated. If no user's first input information is received within the first preset time period, it is assumed that no command was generated based on the user's input information within the first preset time period. Then, based on the preset calling order, the calling command of the second preset node corresponding to the start command is generated. The preset calling order and the first preset time period can be flexibly set based on the actual situation.

[0041] In some embodiments, during the process of displaying the configuration file corresponding to the node in the roaming system, the process of triggering the call instruction according to the preset strategy can be implemented as follows: receiving the call completion instruction of the first node and not receiving the first input information within the second preset time period, generating the call instruction of the third preset node corresponding to the first node based on the preset call order.

[0042] The second preset time period can be flexibly set based on the actual situation.

[0043] The completion call instruction can be generated if the configuration file content corresponding to the first node has been displayed and no user input information has been received within a second preset time period after the display is completed. If no instruction based on user input information has been generated within the second preset time period, then a call instruction for the third preset node corresponding to the first node is generated based on the preset call order.

[0044] In some embodiments, generating a call instruction for roaming video includes: receiving user input information; if the input information is determined to be the first input information by a large language model, determining the output tool corresponding to the input information as a scene switching tool; determining the target node corresponding to the first input information by the large language model; and generating a call instruction corresponding to the target node by the scene switching tool.

[0045] As mentioned above, the invocation command can be generated based on the received user input information, but the input information is not necessarily used to generate the invocation command, and the input information needs to be identified.

[0046] In some embodiments, before receiving user input, a wake-up message for the input can be received first. If a wake-up word or gesture is received, preparation can be made to receive user input. In some embodiments, user input information can be received, and the type of input information can be determined through a large language model. If the input information is determined to be the first input information through the large language model, that is, if the input information includes call information for the target node, then the output tool corresponding to the input information is determined to be a scene switching tool, and the target node corresponding to the first input information is determined; the call instruction corresponding to the target node is generated through the scene switching tool.

[0047] Among them, the large language model can determine the content of the input information in order to determine whether it includes calling information for the target node, and determine the tool to be applied based on the content of the input information.

[0048] In some embodiments, the process of applying the tool can be implemented as follows: determining the target application tool based on the content of the input information, and invoking the target application tool based on the functional description and parameter model of the target application tool.

[0049] Furthermore, the scene switching tool can be used to generate the calling instructions corresponding to the target node.

[0050] In some embodiments, by pre-configuring a list of hot words containing scene-specific nouns, the determination result is biased during the process of determining input information through a large language model, thereby dynamically increasing the score of the lexical sequence corresponding to the hot words and improving the recognition accuracy of the specific words.

[0051] In some embodiments, before determining that the input information is the first input information through the large language model and before determining that the output tool corresponding to the input information is the scene switching tool, the method further includes: if the input information is determined to be text information, then inputting the text information into the large language model; if the input information is determined not to be text information, then determining the information type corresponding to the input information; determining the text information corresponding to the input information based on the text conversion tool corresponding to the information type; and inputting the text information into the large language model.

[0052] In some embodiments, the input information can be text information or other information. Large language models can generally only recognize text information. Therefore, if the input information is identified as text information, the large language model directly recognizes the text information. If the input information is not text information, the information type corresponding to the input information is determined. Based on the information type, the text information corresponding to the input information is determined by the text-to-text tool, and the converted text information is input into the large language model for recognition.

[0053] In some embodiments, if the input information is voice input, the roaming system is generally equipped with a microphone, which can be used to acquire the voice input. After determining that the user's input information is voice information, the input information is input to the speech-to-text tool so that the speech-to-text tool can input the text information corresponding to the input information.

[0054] In some embodiments, the method includes: receiving input information from a user; if the input information is determined to be second input information by a large language model, then determining the output tool of the input information as a voice response tool; determining the response information of the second input information by the large language model, the second input information including question information for the roaming video; inputting the response information into the voice response tool; and obtaining the target response information output by the voice response tool, the target response information being a response information in voice form.

[0055] In some embodiments, user input information can be received at any time after roaming is initiated. If it is determined that the input information includes questions about the roaming video, i.e., the input information is determined to be the second input information, the output tool for the input information is determined to be a voice response tool. The response information for the second input information is determined through a large language model, and the response information is input into the voice response tool. The target response information output by the voice response tool is then obtained.

[0056] In some embodiments, after providing the narrator's image information corresponding to the target node, the method further includes: Upon receiving a user's request to switch the guide's image information, the guide's image information is switched. The input method for the switching information is any of the following: user gesture, keyboard input, or gamepad input.

[0057] There are multiple narrator images corresponding to the target node. If, after displaying one narrator image, a user requests a switch for the narrator image, the narrator image information is switched.

[0058] In some embodiments, the target configuration file also includes a 3D roaming scene corresponding to the target node. After displaying the panoramic video data corresponding to the target node, the narrator's image information corresponding to the target node, the narration information of the target node, and the target 3D model corresponding to the target node, the method further includes: receiving the user's entry command for the 3D digital environment and displaying the 3D roaming scene corresponding to the target node based on the user's location.

[0059] In some embodiments, after displaying the panoramic video data corresponding to the target node, the narrator's image information corresponding to the target node, the narration information of the target node, and the target 3D model corresponding to the target node, if a user's entry command for the 3D digital environment is received, the 3D roaming scene corresponding to the target node is displayed based on the user's location.

[0060] Users will enter a 3D reconstructed landscape scene. The system uses depth cameras and other methods to pinpoint the user's location, updating the camera position in real-time to enable interactive walking within the 3D reconstructed scene. The content of the 3D walkthrough module is either a mesh model or generated based on 3D Gaussian Splatting (3DGS).

[0061] In some embodiments, the call completion instruction of the first node can be generated after the configuration file content corresponding to the first node has been displayed, or it can be generated after the user enters and exits the 3D roaming scene.

[0062] In some embodiments, the above-described method for generating roaming videos and the process of displaying roaming videos are illustrated through the following specific examples.

[0063] Assume the roaming video is a landscape scene roaming video, and the roaming video consists of 5 landscape nodes, referred to as node 1 to node 5, and the corresponding landscapes are landscape 1 to landscape 5 respectively. The preset calling order is node 1 to node 5, and the travel route of each node is the travel route from the previous node to the current node in the calling order.

[0064] Assuming that after the roaming system is started, the user is at the entrance of the landscape, if the system receives the start command and does not receive the user's first input information within the first preset time period, the system generates a call command including node 1 based on the preset call order and displays the travel route from the entrance to node 1.

[0065] Display the panoramic video data corresponding to node 1, the image information of the narrator corresponding to node 1, the narration information of node 1, and the target 3D model corresponding to node 1. The target 3D model can be a 3D model of the landmark building corresponding to node 1.

[0066] After a period of time, the system receives a request from the user to switch the guide's image information and then switches the guide's image information accordingly.

[0067] After displaying the configuration file content corresponding to node 1, the system receives the user's entry command for the 3D digital environment and displays the 3D roaming scene corresponding to the target node based on the user's location.

[0068] After the 3D walkthrough ends, if no initial input from the user is received, a call command including node 2 is generated, and the contents of the configuration file corresponding to node 2 are displayed in the manner described above.

[0069] In some embodiments, if user input is received at any time after the roaming system is started, a corresponding calling instruction or a response to the input can be generated based on the input.

[0070] Assuming the user is at the doorway and issues the voice command "Take me to node 3," a call command including node 3 is generated, and the route from node 2 to node 3 is played. The configuration file content corresponding to node 3 is further displayed.

[0071] At any point after the aforementioned roaming system is activated, it receives a user's question. The large language model identifies the question, understands the user's intent as an information query, selects the "voice answer" tool, and generates text information about the question. This text information is then used by a digital human guide to answer the user's question with voice and realistic lip-sync animation, achieving an immersive interactive Q&A experience.

[0072] In some embodiments, this application also provides a flowchart of a roaming video generation method, such as... Figure 2As shown: This includes the landscape display section and the user interaction section.

[0073] The landscape display section is as follows: First, the system receives the start command of the roaming system. If no user input information is received within the first preset time period, it generates the call command for the second preset node corresponding to the start command based on the preset call order. It then plays the travel route corresponding to the second preset node and displays the panoramic video data of the second preset node, the image information of the target narrator corresponding to the target node, the narration information of the target node, and the target 3D model corresponding to the target node.

[0074] If a user requests a change in the guide's image information, the guide's image information will be changed.

[0075] If a user's command to enter a 3D roaming scene is received, the 3D roaming scene corresponding to the target node will be displayed based on the user's location.

[0076] After displaying the 3D roaming scene corresponding to the target node, the next node is retrieved and the above landscape display process is executed.

[0077] User interaction section: At any moment when the roaming system is activated, it receives a wake-up message, such as a wake-up word or a wake-up gesture, and receives user input. If the input is text, it is directly input into the large language model. If the input is non-text, it is converted into text using a corresponding conversion tool and then input into the large language model. For example, if the input is speech, the corresponding text is generated using a speech-to-text tool.

[0078] Further determine the content of the input information. If the input information includes a call to a target node, generate a call instruction for the target node through the scene switching tool and display the landscape corresponding to the target node. If the input information includes a question for the roaming video, output the corresponding voice answer information through the voice answer tool.

[0079] This application also provides a roaming video generation apparatus, which is used to execute the roaming video generation method provided in any of the above embodiments. For example... Figure 3 As shown, the device includes a generation module 301, a calling module 302, a playback module 303, and a display module 304.

[0080] The generation module 301 is used to generate a calling instruction for the roaming video. The calling instruction includes the target node to be called. The roaming video includes multiple nodes, and each node corresponds to a configuration file. The calling module 302 is used to call the target configuration file corresponding to the target node based on the calling instruction. The target configuration file includes the travel route from the first preset node to the target node, the panoramic video data corresponding to the target node, the narrator image information corresponding to the target node, the narration information of the target node, and the target 3D model corresponding to the target node. The playback module 303 is used to play the travel route from the first preset node to the target node; The display module 304 is used to display the panoramic video data corresponding to the target node, the image information of the narrator corresponding to the target node, the narration information of the target node, and the target 3D model corresponding to the target node.

[0081] This application proposes a roaming video generation device, comprising: generating a calling instruction for the roaming video, the calling instruction including a target node to be called, the roaming video including multiple nodes, each node corresponding to a configuration file; calling the target configuration file corresponding to the target node based on the calling instruction, the target configuration file including a travel route from a first preset node to the target node, panoramic video data corresponding to the target node, narrator image information corresponding to the target node, narration information of the target node, and a target 3D model corresponding to the target node; playing the travel route from the first preset node to the target node; and displaying the panoramic video data corresponding to the target node, the narrator image information corresponding to the target node, the narration information of the target node, and the target 3D model corresponding to the target node. This application embodiment, by configuring corresponding configuration files for multiple nodes of the roaming video respectively, allows for adjustment of the display order of nodes based on user needs, and only requires changing the configuration file corresponding to the roaming video to change the displayed roaming video, making the operation simple.

[0082] In some embodiments, the generation module 301 is specifically used for: If the system receives a roaming system activation command and does not receive the user's first input information within a first preset time period, it generates a second preset node activation command corresponding to the activation command based on a preset activation order. The first input information includes activation information for the target node. Alternatively, if the call completion instruction of the first node is received and the first input information is not received within a second preset time period, a call instruction for a third preset node corresponding to the first node is generated based on the preset call order, wherein the first node is any of the plurality of nodes.

[0083] In some embodiments, the generation module 301 is further specifically used for: Receive user input information; If the input information is determined to be the first input information through a large language model, the output tool corresponding to the input information is determined to be a scene switching tool; The target node corresponding to the first input information is determined using a large language model; The scene switching tool generates the calling instruction corresponding to the target node.

[0084] In some embodiments, the generation module 301 is further specifically used for: If the input information is determined to be text information, then the text information is input into the large language model; If it is determined that the input information is not text information, then the information type corresponding to the input information is determined; The text information corresponding to the input information is determined based on the text conversion tool corresponding to the information type; The text information is input into the large language model.

[0085] In some embodiments, the above-described apparatus further includes: The receiving module is used to receive user input information. The determination module is used for: If the input information is determined to be the second input information through a large language model, then the output tool for the input information is determined to be a voice response tool; The response information for the second input information is determined using a large language model, and the second input information includes question information for the roaming video. An input module is used to input the answer information into the voice answer tool; The acquisition module is used to acquire the target answer information output by the voice answer tool, wherein the target answer information is answer information in voice form.

[0086] In some embodiments, the receiving module is further configured to: Upon receiving a user's request to switch the narrator's image information, the narrator's image information is switched. The input method for the switching information is any of the following: user gesture, keyboard input, or gamepad input.

[0087] In some embodiments, the target configuration file further includes a 3D roaming scene corresponding to the target node, and the display module 304 is further used for: Receives user's entry command for a 3D digital environment and displays a 3D roaming scene corresponding to the target node based on the user's location.

[0088] This application also provides an electronic device for performing the above-described roaming video generation method. Please refer to... Figure 4It illustrates a schematic diagram of an electronic device provided by some embodiments of this application. For example... Figure 4 As shown, the electronic device 7 includes: a processor 700, a memory 701, a bus 702 and a communication interface 703. The processor 700, the communication interface 703 and the memory 701 are connected through the bus 702. The memory 701 stores a computer program that can run on the processor 700. When the processor 700 runs the computer program, it executes the roaming video generation method provided in any of the foregoing embodiments of this application.

[0089] The memory 701 may include high-speed random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Communication between the device network element and at least one other network element is achieved through at least one communication interface 703 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc.

[0090] Bus 702 can be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. The memory 701 is used to store programs. After receiving execution instructions, the processor 700 executes the program. The roaming video generation method disclosed in any of the aforementioned embodiments of this application can be applied to the processor 700, or implemented by the processor 700.

[0091] The processor 700 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of the processor 700 or by instructions in software form. The processor 700 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software units in the decoding processor. The software units may reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 701. Processor 700 reads the information in memory 701 and, in conjunction with its hardware, completes the steps of the above method.

[0092] The electronic device provided in this application embodiment and the roaming video generation method provided in this application embodiment are based on the same inventive concept and have the same beneficial effects as the methods they adopt, operate or implement.

[0093] This application also provides a computer-readable storage medium corresponding to the roaming video generation method provided in the foregoing embodiments. Please refer to... Figure 5 The computer-readable storage medium shown is an optical disc 30, on which a computer program (i.e., a program product) is stored. When the computer program is run by a processor, it executes the roaming video generation method provided in any of the aforementioned embodiments.

[0094] It should be noted that examples of computer-readable storage media may also include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical and magnetic storage media, which will not be elaborated here.

[0095] The computer-readable storage medium provided in the above embodiments of this application and the roaming video generation method provided in the embodiments of this application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the applications stored therein.

[0096] It should be noted that: Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of this application may be practiced without these specific details. In some instances, well-known structures and techniques have not been shown in detail so as not to obscure the understanding of this specification.

[0097] Furthermore, those skilled in the art will understand that although some embodiments herein include certain features included in other embodiments but not others, combinations of features from different embodiments are intended to be within the scope of this application and form different embodiments. For example, in the following claims, any of the claimed embodiments can be used in any combination.

[0098] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for generating roaming videos, characterized in that, include: Generate a calling instruction for the roaming video, the calling instruction including the target node to be called, the roaming video including multiple nodes, each node corresponding to a configuration file; The target configuration file corresponding to the target node is invoked based on the invocation instruction. The target configuration file includes the travel route from the first preset node to the target node, the panoramic video data corresponding to the target node, the image information of multiple narrators corresponding to the target node, the narration information of the target node, and the target 3D model corresponding to the target node. Play the travel route from the first preset node to the target node; The system displays the panoramic video data corresponding to the target node, the image information of the target narrator corresponding to the target node, the narration information of the target node, and the target 3D model corresponding to the target node.

2. The method according to claim 1, characterized in that, The generation of invocation instructions for the roaming video includes: If the system receives a roaming system activation command and does not receive the user's first input information within a first preset time period, it generates a second preset node activation command corresponding to the activation command based on a preset activation order. The first input information includes activation information for the target node. Alternatively, if the call completion instruction of the first node is received and the first input information is not received within a second preset time period, a call instruction for a third preset node corresponding to the first node is generated based on the preset call order, wherein the first node is any of the plurality of nodes.

3. The method according to claim 2, characterized in that, The generation of invocation instructions for the roaming video includes: Receive user input information; If the input information is determined to be the first input information through a large language model, the output tool corresponding to the input information is determined to be a scene switching tool; The target node corresponding to the first input information is determined using a large language model; The scene switching tool generates the calling instruction corresponding to the target node.

4. The method according to claim 3, characterized in that, Before determining that the input information is the first input information through a large language model, and before determining that the output tool corresponding to the input information is a scene switching tool, the method further includes: If the input information is determined to be text information, then the text information is input into the large language model; If it is determined that the input information is not text information, then the information type corresponding to the input information is determined; The text information corresponding to the input information is determined based on the text conversion tool corresponding to the information type; The text information is input into the large language model.

5. The method according to claim 1, characterized in that, After playing the travel route from the first preset node to the target node, the method includes: Receive user input information; If the input information is determined to be the second input information through a large language model, then the output tool for the input information is determined to be a voice response tool; The response information for the second input information is determined using a large language model, and the second input information includes question information for the roaming video. Input the answer information into the voice answer tool; Obtain the target answer information output by the voice answer tool, wherein the target answer information is in the form of voice answer information.

6. The method according to claim 1, characterized in that, After displaying the narrator's image information corresponding to the target node, the method further includes: Upon receiving a user's request to switch the narrator's image information, the narrator's image information is switched. The input method for the switching information is any of the following: user gesture, keyboard input, or gamepad input.

7. The method according to claim 1, characterized in that, The target configuration file also includes the 3D roaming scene corresponding to the target node. After displaying the panoramic video data corresponding to the target node, the narrator's image information corresponding to the target node, the narration information of the target node, and the target 3D model corresponding to the target node, the method further includes: Receives user's entry command for a 3D digital environment and displays a 3D roaming scene corresponding to the target node based on the user's location.

8. A roaming video generation device, characterized in that, include: The generation module is used to generate a calling instruction for the roaming video. The calling instruction includes the target node to be called. The roaming video includes multiple nodes, and each node corresponds to a configuration file. The calling module is used to call the target configuration file corresponding to the target node based on the calling instruction. The target configuration file includes the travel route from the first preset node to the target node, the panoramic video data corresponding to the target node, the narrator image information corresponding to the target node, the narration information of the target node, and the target 3D model corresponding to the target node. The playback module is used to play the travel route from the first preset node to the target node; The display module is used to display the panoramic video data corresponding to the target node, the image information of the narrator corresponding to the target node, the narration information of the target node, and the target 3D model corresponding to the target node.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the method as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by a processor to implement the method as described in any one of claims 1-7.