Immersive AI Scenic Area Data Output Method and System

By generating personalized video clips and dynamically adjusting screen resources, the issues of flexibility and interactivity in multimedia display methods in scenic areas have been resolved, thereby enhancing the immersive experience for tourists and improving resource utilization efficiency.

CN120803271BActive Publication Date: 2026-07-17NANJING MOCHOU INTELLIGENT INFORMATION TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING MOCHOU INTELLIGENT INFORMATION TECH CO LTD
Filing Date
2025-07-16
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

The existing multimedia display methods in scenic spots cannot be customized according to tourists' personalized tour routes and interests, resulting in tourists being unable to effectively interact with the display content. Resource planning lacks flexibility, leading to idle and wasted resources and failing to meet the viewing needs during peak periods, thus limiting the immersive experience.

Method used

By generating video segments based on user location and interests, combined with dynamic adjustments to the immersive data acquisition area and output screen, personalized data output segments are generated and interactively displayed on idle screens, leveraging AI modules to enhance interactivity.

Benefits of technology

It enables dynamic adjustment of screen resources based on tourist needs, enhancing immersive experience and interactivity, making full use of screen resources, and increasing user engagement and experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120803271B_ABST
    Figure CN120803271B_ABST
Patent Text Reader

Abstract

This invention provides an immersive AI-powered scenic area data output method and system. The method includes: determining a first video segment corresponding to each user terminal that meets the immersive data output conditions based on immersive data acquisition areas at each first location; upon any user terminal arriving at any corresponding second location's immersive data output area, acquiring and decomposing the immersive output points of the first video segment based on the user points, generating data output segments, and determining the number of output screens corresponding to each data output segment; and upon responding to a situation where the number of output screens with idle attributes located in the immersive data output area is greater than or equal to the total number of output screens, interactively displaying each data output segment to the user terminal based on each of the data output screens. This invention at least improves the data output effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to data processing technology, and more particularly to an immersive AI-powered scenic area data output method and system. Background Technology

[0002] As people's living standards improve, tourism has become an important form of leisure and entertainment. In various scenic areas, tourists are no longer satisfied with superficial sightseeing; instead, they expect to use advanced technology to deeply experience the unique features of the area and gain unique, interactive enjoyment. The demand for immersive travel experiences is becoming increasingly strong. Currently, scenic spots generally adopt traditional multimedia display methods, such as playing standardized promotional videos in a fixed area or setting up fixed screens to display scenic landscapes, history, and culture. However, these methods cannot be customized according to tourists' individual tour routes and interests, making it impossible for tourists to effectively interact with the displayed content. Furthermore, the planning of screen resources within scenic spots lacks flexibility, making it difficult to dynamically adjust usage based on tourists' real-time needs and on-site conditions. This results in both resource idleness and waste, and an inability to meet the viewing needs of tourists during peak periods, greatly limiting the immersive experience for tourists. Summary of the Invention

[0003] Based on the above problems, the present invention is proposed to provide an immersive AI scenic spot data output method and system that overcomes or at least partially solves the above problems.

[0004] According to one aspect of the present invention, an immersive AI scenic spot data output method is provided, comprising the following steps: Based on the immersive data acquisition area of ​​each first position, the first video segment corresponding to each user terminal that meets the immersive data output conditions is determined. Upon responding to any user terminal reaching any corresponding second location of the immersive data output area, the immersive output point of the first video segment is acquired and decomposed based on the user point, a data output segment is generated, and the number of output screens for the corresponding data output segment is determined. If the number of output screens with idle attributes located in the immersive data output area is greater than or equal to the number of output screens, the data output segments are interactively displayed to the user based on each of the data output screens.

[0005] Optionally, in the method according to the present invention, determining each first video segment corresponding to each user terminal that satisfies the immersive data output conditions based on the immersive data acquisition area of ​​each first location includes: If the location information of any user terminal overlaps with the location of any immersive data acquisition area, determine the overlapping time period corresponding to the overlapping relationship; If the overlap duration of the corresponding overlapping time period is greater than the preset duration, it is determined that the user terminal meets the immersive data output conditions; Based on the orientation acquisition unit corresponding to the immersive data acquisition area, the orientation acquisition video of the corresponding overlapping time period is obtained, and the orientation acquisition video is recognized based on the video recognition model to obtain the viewing orientation of the corresponding user terminal and the viewing time of each viewing orientation. The viewing orientation corresponding to the maximum viewing time is determined as the user's main orientation at the corresponding user terminal, and the scenic area acquisition unit corresponding to the user's main orientation acquires the scenic area acquisition video of the corresponding overlapping time period. The video footage collected from the scenic area is identified as the first video segment corresponding to the user's device.

[0006] Optionally, in the method according to the present invention, after responding to any of the user terminals reaching the immersive data output area corresponding to any second position, the immersive output points of the first video segment are acquired and decomposed based on the user points to generate each data output segment, and the number of output screens corresponding to each data output segment is determined, including: In response to any user terminal arriving at any corresponding second location of the immersive data output area, retrieve each first video segment corresponding to the user terminal; In response to user interaction with any first video segment, retrieve immersive experience images of the immersive data acquisition area corresponding to the first video segment, wherein the immersive experience images include each immersive experience point in the corresponding immersive data acquisition area; In response to user interaction with any immersive experience point based on the immersive experience image, the immersive experience point is determined as the immersive output point. Based on the user point, data output segments corresponding to the user are generated, and the number of output screens corresponding to each data output segment is determined.

[0007] Optionally, in the method according to the present invention, generating data output segments corresponding to user terminals based on user location decomposition and determining the number of output screens corresponding to each data output segment includes: Based on the acquisition units of each point of the corresponding immersive output point, the acquisition video of each point during the corresponding overlapping time period is obtained, and the target action attribute of each point acquisition video is determined, wherein the target action attribute includes target action and no target action. Based on the preset output time, the video of each point that is captured with the corresponding target attribute of having target action is decomposed into points to obtain each data output segment; The number of data outputs for each corresponding data output segment is determined as the number of output screens.

[0008] Optionally, in the method according to the present invention, the target action attribute of the video captured at each point is determined, wherein the target action attribute includes target action and no target action, including: Identify the individual image frames that make up the video captured at each location, and perform image recognition on each image frame. Based on the recognition results, the captured image frames of each animal area corresponding to the same location are aggregated into the same animal image group, and the video of the location is updated based on the captured image frames located in the animal image group. If the duration of the video captured at any point after the corresponding video update exceeds a preset time threshold, the target attribute of the video captured at that point is determined to have a target action. Conversely, the target attribute of the video captured at that location is determined to be a targetless action.

[0009] Optionally, in the method according to the present invention, the video footage collected at each point corresponding to the target attribute of having a target action is decomposed based on a preset output time to obtain each data output segment, including: Retrieve each captured image frame from the video of each point where the target attribute is a target action, and determine the center point of each area of ​​each animal area that indicates each animal to be viewed, located in each captured image frame. Based on the video captured at each location, the center points of each region of the same animal are divided into the same location group according to the images captured in adjacent frames. Each image coordinate system is established with the center point of each acquired image frame as the origin, and the coordinate points of each region corresponding to the center point of each region are determined based on each image coordinate system. The distances between the coordinates of each area corresponding to the same location group are determined based on the distances between the coordinates of each area. If the distance to the response area is greater than the first preset distance, the acquired image frames corresponding to the distance to the area will be aggregated into a moving image group corresponding to the point acquisition video of each acquired image frame. The number of image frames corresponding to the same moving image group is determined based on each acquired image frame located in the same moving image group, and the number of image frames is compared with the number of frames determined based on the preset output time. If the number of response image frames is greater than or equal to a preset number of frames, video fusion is performed on each of the acquired image frames based on the preset number of frames, and point decomposition is performed on each of the resulting data output segments.

[0010] Optionally, in the method according to the invention, the method further includes: If the number of response image frames is less than the preset number of frames, determine the number of animal regions in each captured image frame in each captured video frame, excluding the corresponding moving image group. If the number of any region exceeds the preset number, the captured image frame corresponding to that region is added to the moving image group so that the captured image frame is connected to the last captured image frame in the moving image group, thus obtaining a moving image group with updated frame count. If the number of image frames in a moving image group that responds to a frame update is less than a preset number of frames, the acquired image frames whose corresponding moving distance is less than a first preset distance but greater than a second preset distance are added to the moving image group in chronological order until the number of image frames in the corresponding moving image group is greater than or equal to the preset number of frames.

[0011] Optionally, in the method according to the present invention, point decomposition is performed on each of the obtained data output segments, including: Determine the percentage of each image corresponding to each animal region in the output image frames that make up each data output segment, indicating each viewer animal; In response to any image proportion being greater than a first preset proportion and less than a second preset proportion, the animal region corresponding to the image proportion is determined as the first target region, and target extraction is performed on the first target region; The extracted first target region is magnified based on a first magnification factor. Since the magnified first target region overlaps with any animal region, the first target region is reduced in image based on a first reduction factor corresponding to the first magnification factor. If any image proportion is less than or equal to a second preset proportion, the animal region corresponding to the image proportion is determined as the second target region, and target extraction is performed on the second target region; The extracted second target region is magnified based on a second magnification factor, wherein the second magnification factor is greater than the first magnification factor; If the magnified second target region overlaps with any animal region, the second target region is image-reduced based on the corresponding second magnification factor and the second target region is determined as the first target region.

[0012] Optionally, in the method according to the present invention, interactively displaying each data output segment to the user terminal based on each of the data output screens includes: Each data output segment is sent to its respective output display slot on each data output screen that has an idle attribute; In response to the user completing the wearing of any interactive unit, the data output segments are output and displayed, and the voice broadcasting unit based on the immersive data output area broadcasts the animal broadcasting templates in the immersive data collection area corresponding to each data output segment. In response to user interaction with any animal limb corresponding to any animal in any data output segment via the interaction unit, the AI ​​module generates an interactive action corresponding to the animal limb with a preset number of interaction frames and displays it to the user. The voice broadcasting unit then broadcasts the interactive broadcasting template corresponding to the interactive action. Based on a preset number of interactive frames, the acquired image frames located at the end of the data output segment are removed.

[0013] According to another aspect of the present invention, an immersive AI scenic area data output system is provided, comprising: The determination module is configured to determine the first video segment corresponding to each user terminal that meets the immersive data output conditions based on the immersive data acquisition area of ​​each first position. The decomposition module is configured to, in response to any user terminal reaching any corresponding second position of the immersive data output area, acquire the immersive output point of the first video segment and decompose it based on the user point to generate a data output segment, and determine the number of output screens for the corresponding data output segment. The interaction module is configured to respond to a situation where the number of output screens with idle attributes located in the immersive data output area is greater than or equal to the number of output screens, and to interactively display each data output segment to the user terminal based on each of the data output screens.

[0014] According to the present invention, firstly, the server can accurately determine the first video segment corresponding to the user terminal based on the viewing situation of the user terminal in the immersive data acquisition area at the first location, providing personalized materials for subsequent immersive experiences. Secondly, after the user terminal reaches the immersive data output area at the second location, the server acquires the immersive output points of the first video segment and decomposes them based on the user's points to generate data output segments, enabling subsequent users to view the animals in an immersive way, enhancing the sense of immersion. Simultaneously, the server determines the number of output screens for the corresponding data output segments, which helps to rationally plan screen resources. Finally, when the number of data output screens with idle attributes in the immersive data output area meets the requirements, the server interactively displays each data output segment to the user terminal based on each data output screen, fully utilizing screen resources and providing users with multi-dimensional, interactive display effects, improving user participation and experience. Attached Figure Description

[0015] Figure 1 A flowchart of an immersive AI scenic area data output method according to an embodiment of the present invention is shown; Figure 2 A schematic diagram of immersive experience points according to an embodiment of the present invention is shown; Figure 3 A schematic diagram of the area distance according to an embodiment of the present invention is shown; Figure 4 A structural block diagram of an immersive AI scenic area data output system according to another embodiment of the present invention is shown. Detailed Implementation

[0016] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0017] To address the problems existing in the aforementioned background technology, the inventors proposed the solution of this invention. One embodiment of this invention provides an immersive AI scenic area data output method, which can be executed in a computing device.

[0018] Figure 1 A flowchart of an immersive AI scenic area data output method according to an embodiment of the present invention is shown, the method being adapted to be executed in a computing device.

[0019] like Figure 1 As shown, the immersive AI scenic area data output method proposed in this embodiment begins with step S102, which includes the following: Based on the immersive data acquisition area of ​​each first position, the first video segment corresponding to each user terminal that meets the immersive data output conditions is determined.

[0020] For example, in this embodiment, the user terminal can be understood as a tourist wearing a positioning wristband. There are multiple immersive data collection areas set up in the scenic area, each located at a first position. The immersive data collection areas can be understood as different animal enclosures in a zoo. In this embodiment, each immersive data acquisition area not only contains different animals for viewing, but also determines the first video segment corresponding to the user terminal based on the viewing situation of the user terminal in the immersive data acquisition area that meets the immersive data output conditions. The generation of the first video segment facilitates the subsequent generation of various data output segments for the user to have an immersive experience by the server.

[0021] Furthermore, the aforementioned "determining each first video segment corresponding to each user terminal that meets the immersive data output conditions based on the immersive data acquisition area of ​​each first location" also includes the following steps: If the location information of any user terminal overlaps with the location of any immersive data acquisition area, determine the overlapping time period corresponding to the overlapping relationship; If the overlap duration of the corresponding overlapping time period is greater than the preset duration, it is determined that the user terminal meets the immersive data output conditions; Based on the orientation acquisition unit corresponding to the immersive data acquisition area, the orientation acquisition video of the corresponding overlapping time period is obtained, and the orientation acquisition video is recognized based on the video recognition model to obtain the viewing orientation of the corresponding user terminal and the viewing time of each viewing orientation. The viewing orientation corresponding to the maximum viewing time is determined as the user's main orientation at the corresponding user terminal, and the scenic area acquisition unit corresponding to the user's main orientation acquires the scenic area acquisition video of the corresponding overlapping time period. The video footage collected from the scenic area is identified as the first video segment corresponding to the user's device.

[0022] For example, in this embodiment, the server will acquire the location information of the user terminal in real time. When the location information of any user terminal overlaps with the location of any immersive data collection area, it indicates that the user terminal has entered the immersive data collection area. Next, the server will determine the duration of the user's time in the immersive data collection area, that is, the overlap duration of the overlapping period when the location information and the area location overlap. When the overlap duration is longer than the preset duration, it means that the user's time in the immersive data collection area is relatively long. In other words, it can be preliminarily determined that the user is more interested in the animals being observed in the immersive data collection area. Therefore, the server will determine that the user meets the immersive data output conditions. In order to make the subsequent data output segments more suitable for the user's viewing experience, the server will first determine which direction of the animals the user is more interested in when watching the animals using the following method. First, the server will acquire the orientation acquisition video of the corresponding overlapping time period through the orientation acquisition unit set in the immersive data acquisition area. Then, the video recognition model will perform video recognition on the orientation acquisition video to obtain the viewing orientation of the corresponding user at the corresponding overlapping time period and the viewing time of each viewing orientation. Next, the server will determine the viewing direction corresponding to the maximum viewing time as the user's main orientation. For example, if the viewing directions of the user are east and west, and the viewing times for each viewing direction are 10 minutes and 15 minutes, the server will determine the west viewing direction as the user's main orientation. Finally, the server will obtain the scenic area video of the corresponding overlapping time period according to the scenic area collection unit facing the user, and determine the scenic area video as the first video segment corresponding to the user, so that the user can quickly select the experience scene they want to immerse themselves in based on each first video segment.

[0023] Step S104 includes the following: Upon responding to any user terminal reaching any corresponding second location of the immersive data output area, the immersive output point of the first video segment is acquired and decomposed based on the user point to generate a data output segment, and the number of output screens for the corresponding data output segment is determined.

[0024] For example, in this embodiment, multiple immersive data output areas will be set up in the scenic area, which are areas for subsequent users to have an immersive experience. When a user arrives at any immersive data output area corresponding to the second position, the server will acquire the immersive output point of the first video segment and decompose it based on the user's point to generate a data output segment. Since the location of the immersive output point may be close to the animals being viewed, the subsequent users can enjoy the animals in an immersive way based on the data output segment. There may be multiple data output screens in the immersive data output area, and the number of output screens for different user terminals is different. Therefore, the server will determine the number of output screens for the corresponding data output segment to facilitate subsequent screen planning for each data output screen, so that each user terminal can quickly and efficiently enjoy the immersive experience.

[0025] Furthermore, the aforementioned "responding to any user terminal reaching any corresponding second location of the immersive data output area, acquiring the immersive output point of the first video segment and decomposing it based on the user point to generate each data output segment, and determining the number of output screens corresponding to each data output segment" also includes the following steps: In response to any user terminal arriving at any corresponding second location of the immersive data output area, retrieve each first video segment corresponding to the user terminal; In response to user interaction with any first video segment, retrieve immersive experience images of the immersive data acquisition area corresponding to the first video segment, wherein the immersive experience images include each immersive experience point in the corresponding immersive data acquisition area; In response to user interaction with any immersive experience point based on the immersive experience image, the immersive experience point is determined as the immersive output point. Based on the user point, data output segments corresponding to the user are generated, and the number of output screens corresponding to each data output segment is determined.

[0026] For example, in this embodiment, when any user terminal arrives at any immersive data output area corresponding to the second position, it can be understood that the user terminal wants to have an immersive experience. At this time, the server will retrieve each first video segment corresponding to the user terminal, so that the user terminal can select the specific scenario for the immersive experience by viewing each first video segment. When a user interacts with any of the first video segments, the server retrieves the immersive experience image of the immersive data acquisition area corresponding to that first video segment. The immersive experience image can be understood as a top view of the area corresponding to the animal viewing area within the immersive data acquisition area, and it contains various immersive experience points, such as... Figure 2 As shown; When a user interacts with any immersive experience point in the immersive experience image, it means that the user wants to experience the immersive experience from the perspective of that point. Therefore, the server will determine that immersive experience point as the immersive output point, and then decompose the user point to generate the corresponding data output segments for the user. The server will also determine the number of output screens for each data output segment, so as to make reasonable arrangements for the user's immersive experience based on the number of output screens.

[0027] Furthermore, the aforementioned "generating data output segments for corresponding user terminals based on user location decomposition, and determining the number of output screens for each data output segment" also includes the following steps: Based on the acquisition units of each point of the corresponding immersive output point, the acquisition video of each point during the corresponding overlapping time period is obtained, and the target action attribute of each point acquisition video is determined, wherein the target action attribute includes target action and no target action. Based on the preset output time, the video of each point that is captured with the corresponding target attribute of having target action is decomposed into points to obtain each data output segment; The number of data outputs for each corresponding data output segment is determined as the number of output screens.

[0028] For example, in this embodiment, in order to provide users with a more realistic and rich immersive experience through the data output segments generated for the immersive output points, multiple acquisition units will be set up at each immersive output point to facilitate the comprehensive video acquisition of the animals in the immersive acquisition area based on the immersive output points. To ensure that the video scene of the subsequently generated data output segment is consistent with the first video segment on the user's end, the server will obtain the video captured at each point of the corresponding overlapping time period according to the acquisition unit of each point of the corresponding immersive output point. If there are no animals to watch or only animals to watch briefly in the data output segment generated later, the immersive experience will be greatly reduced when the user is having an immersive experience, and unnecessary screen consumption will also increase. Therefore, the server will first determine the target action attributes of the video captured at each point. Specifically, the target action attributes include target actions and no target actions. To improve the efficiency of immersive experiences for various users, the server will decompose the video captured at each point with the corresponding target attribute of targeted action according to the preset output time, thereby obtaining each data output segment. The preset output time can be preset by the management terminal according to the actual situation. Finally, the server will determine the number of data outputs for each data output segment as the number of output screens, which will facilitate subsequent planning of each data output screen so that users can quickly have an immersive experience.

[0029] Furthermore, the aforementioned "determining the target action attributes of the video collected at each location, wherein the target action attributes include target actions and non-target actions" also includes the following steps: Identify the individual image frames that make up the video captured at each location, and perform image recognition on each image frame. Based on the recognition results, the captured image frames of each animal area corresponding to the same location are aggregated into the same animal image group, and the video of the location is updated based on the captured image frames located in the animal image group. If the duration of the video captured at any point after the corresponding video update exceeds a preset time threshold, the target attribute of the video captured at that point is determined to have a target action. Conversely, the target attribute of the video captured at that location is determined to be a targetless action.

[0030] For example, in this embodiment, the server first determines each captured image frame that makes up the video of each location, and then performs image recognition on each captured image frame to determine whether there is an animal area indicating the animals to be viewed in each captured image frame. Next, the server will aggregate the captured image frames of each animal region that indicate the presence of any animal in the video captured at the same location into the same animal image group based on the recognition results. Then, the server will update the video captured at the location based on each captured image frame in the animal image group, so that the animal region indicating the animal exists in each captured image frame that makes up the updated video captured at the location. If the video duration of the updated location capture video is too short, that is, it cannot provide users with enough immersive experience, which will reduce the user's experience. Therefore, the server will compare the video duration of the corresponding location capture video with a preset time threshold. One possible comparison result is that the video duration is greater than the preset time threshold. This comparison result indicates that the video duration of the corresponding point is relatively long, which means that a sufficient immersive experience can be provided to the user based on the video captured at that point. Therefore, the server will determine the target attribute of the video captured at that point as having a target action. Another possible comparison result is that the video duration is less than or equal to the preset time threshold. This comparison result indicates that the video duration of the video collected at the corresponding point is too short, which means that it is impossible to provide a sufficient duration of immersive experience to the user based on the video collected at that point. Therefore, the server will determine the target attribute of the video collected at that point as no target action.

[0031] Furthermore, the aforementioned "decomposing the video footage collected from each point with the corresponding target attribute of having a target action based on a preset output time to obtain each data output segment" also includes the following steps: Retrieve each captured image frame from the video of each point where the target attribute is a target action, and determine the center point of each area of ​​each animal area that indicates each animal to be viewed, located in each captured image frame. Based on the video captured at each location, the center points of each region of the same animal are divided into the same location group according to the images captured in adjacent frames. Each image coordinate system is established with the center point of each acquired image frame as the origin, and the coordinate points of each region corresponding to the center point of each region are determined based on each image coordinate system. The distances between the coordinates of each area corresponding to the same location group are determined based on the distances between the coordinates of each area. If the distance to the response area is greater than the first preset distance, the acquired image frames corresponding to the distance to the area will be aggregated into a moving image group corresponding to the point acquisition video of each acquired image frame. The number of image frames corresponding to the same moving image group is determined based on each acquired image frame located in the same moving image group, and the number of image frames is compared with the number of frames determined based on the preset output time. If the number of response image frames is greater than or equal to a preset number of frames, video fusion is performed on each of the acquired image frames based on the preset number of frames, and point decomposition is performed on each of the resulting data output segments.

[0032] For example, in this embodiment, it can be understood that when users are having an immersive experience, they prefer to see animals in motion rather than animals that are stationary. Therefore, the server will first determine the movement speed of the animals that are being captured in video at each location. First, the server will retrieve each captured image frame of the video from each point where the target attribute is "target action", and then determine the center point of each animal area that indicates each animal to be viewed in each captured image frame. Next, the server will divide the center points of each region of the same animal into the same location group according to the video captured at each location. Then, it will establish each image coordinate system with the center point of each captured image frame as the origin, and determine the coordinate points of each region corresponding to the center point of each region according to each image coordinate system. Next, the server will determine the coordinate distance between the two area coordinate points based on the area coordinates of the center points of the two areas corresponding to the same location group, and will then define each coordinate distance as the area distance for each animal area indicating the same viewing animal, such as... Figure 3 As shown; The greater the area distance, the faster the corresponding animal moves. Therefore, the server will aggregate all captured image frames with an area distance greater than the first preset distance into a moving image group corresponding to the point captured video of each captured image frame. Then, the server will determine the number of image frames for each captured image frame in the same moving image group, and then determine the preset number of frames for the corresponding preset output time. At this time, the server will compare the number of image frames with the preset number of frames. When the number of image frames is greater than or equal to the preset number of frames, it means that the video duration of the video composed of the corresponding number of image frames can meet the duration requirements for the user to have an immersive experience. Therefore, the server will perform video fusion on each captured image frame according to the preset number of frames to obtain the data output segment with the corresponding preset output time. Finally, to enhance the immersive experience for users, the server will further decompose the various data output segments into points.

[0033] Furthermore, the above method also includes the following steps: If the number of response image frames is less than the preset number of frames, determine the number of animal regions in each captured image frame in each captured video frame, excluding the corresponding moving image group. If the number of any region exceeds the preset number, the captured image frame corresponding to that region is added to the moving image group so that the captured image frame is connected to the last captured image frame in the moving image group, thus obtaining a moving image group with updated frame count. If the number of image frames in a moving image group that responds to a frame update is less than a preset number of frames, the acquired image frames whose corresponding moving distance is less than a first preset distance but greater than a second preset distance are added to the moving image group in chronological order until the number of image frames in the corresponding moving image group is greater than or equal to the preset number of frames.

[0034] For example, in this embodiment, when the number of image frames is less than the preset number of frames, it means that the video duration of the video composed of the corresponding number of image frames cannot meet the duration requirement for the user to have an immersive experience. At this time, the server will further filter the remaining captured image frames and add the captured image frames that meet the requirements to the moving image group, so that the number of image frames in the corresponding moving image group is greater than or equal to the preset number of frames. In order to enable users to see more animals during the immersive experience, the server will determine the number of animal areas in each captured image frame, excluding the corresponding moving image group, in the video captured at each location. When the number of regions exceeds the preset number, it indicates that there are too many regions corresponding to the animal regions for the image frames to be captured. Therefore, the server will add the captured image frames corresponding to the number of regions to the moving image group. The captured image frames will be added after the last captured image frame in the moving image group, resulting in a moving image group with updated frame count. This will minimize the removal of captured image frames corresponding to the moving animals when the data output segment determined by the moving image group is removed according to the preset number of interactive frames. If the number of image frames in the corresponding moving image group is still less than the preset number of frames, the server will further add each acquired image frame whose corresponding moving distance is less than the first preset distance but greater than the second preset distance to the moving image group in chronological order until the number of image frames in the corresponding moving image group is greater than or equal to the preset number of frames.

[0035] Furthermore, the aforementioned "decomposition of the obtained data output segments into points" also includes the following steps: Determine the percentage of each image corresponding to each animal region in the output image frames that make up each data output segment, indicating each viewer animal; In response to any image proportion being greater than a first preset proportion and less than a second preset proportion, the animal region corresponding to the image proportion is determined as the first target region, and target extraction is performed on the first target region; The extracted first target region is magnified based on a first magnification factor. Since the magnified first target region overlaps with any animal region, the first target region is reduced in image based on a first reduction factor corresponding to the first magnification factor. If any image proportion is less than or equal to a second preset proportion, the animal region corresponding to the image proportion is determined as the second target region, and target extraction is performed on the second target region; The extracted second target region is magnified based on a second magnification factor, wherein the second magnification factor is greater than the first magnification factor; If the magnified second target region overlaps with any animal region, the second target region is image-reduced based on the corresponding second magnification factor and the second target region is determined as the first target region.

[0036] For example, in this embodiment, since the distance between the viewing animal and the acquisition unit may be different, the proportion of each animal region indicating any viewing animal in each output image frame that makes up each data output segment is different from the proportion of the corresponding image in the output image frame. When the image proportion is too small, it will cause the user to not be able to see the animal clearly when they are immersing themselves in the experience. In order to avoid reducing the user experience, the server will first determine the proportion of each image corresponding to each animal area in the output image frame that indicates any animal in the data output segment. In order to ensure that the proportions of each image in each animal area are relatively close and to facilitate a consistent immersive experience for users, the management can pre-set a first preset proportion and a second preset proportion on the server. One possible scenario is that the image proportion is greater than the first preset proportion and less than the second preset proportion. That is, the animal area corresponding to the image proportion is not very small and only needs to be slightly enlarged. In this case, the server will determine the animal area corresponding to the image proportion as the first target area and extract the target from the first target area. Next, the server will magnify the extracted first target area according to the first magnification factor. In order to ensure that users can see the complete animal areas when they are immersed in the experience, the server will further determine whether the magnified first target area overlaps with other animal areas. If there is an overlap, the first target area will be reduced in size based on the first reduction factor corresponding to the first magnification factor. That is, the image proportion of the reduced first target area is the same as the initial image proportion. Another possible scenario is that the image proportion is less than or equal to the second preset proportion, which means that the animal region corresponding to the image proportion is too small and needs to be greatly enlarged. Therefore, the server will determine the animal region corresponding to the image proportion as the second target region and then extract the target from the second target region. Next, the server will magnify the extracted second target region according to the second magnification factor. It should be noted that the second magnification factor is greater than the first magnification factor. Similarly, the server will further determine whether the magnified second target region overlaps with other animal regions, and if there is an overlap, the second target region will be image-reduced based on the corresponding second magnification factor and the second reduction factor. In order to magnify the image of the second target region as much as possible, the server will determine the second target region as the first target region, and then magnify the image of it based on the first magnification factor in the manner described above.

[0037] Step S106 includes the following: If the number of output screens with idle attributes located in the immersive data output area is greater than or equal to the number of output screens, the data output segments are interactively displayed to the user based on each of the data output screens.

[0038] For example, in this embodiment, there may be multiple data output screens in the immersive data output area. When the number of output screens with idle attributes is greater than or equal to the number of output screens, it means that the immersive data output area has met the screen quantity requirements for providing an immersive experience for the user. At this time, the server will interactively display each data output segment to the user based on each data output screen.

[0039] Furthermore, the aforementioned "interactive display of each data output segment to the user terminal based on each of the data output screens" also includes the following steps: Each data output segment is sent to its respective output display slot on each data output screen that has an idle attribute; In response to the user completing the wearing of any interactive unit, the data output segments are output and displayed, and the voice broadcasting unit based on the immersive data output area broadcasts the animal broadcasting templates in the immersive data collection area corresponding to each data output segment. In response to user interaction with any animal limb corresponding to any animal in any data output segment via the interaction unit, the AI ​​module generates an interactive action corresponding to the animal limb with a preset number of interaction frames and displays it to the user. The voice broadcasting unit then broadcasts the interactive broadcasting template corresponding to the interactive action. Based on a preset number of interactive frames, the acquired image frames located at the end of the data output segment are removed.

[0040] For example, in this embodiment, the server will send each data output segment to the output display slot of each data output screen with idle attribute, and after the user completes wearing any interactive unit, the server will output and display each data output segment to the user, that is, the user begins to immerse in the interaction. Here, the interactive unit can be understood as VR glasses, VR gloves and other interactive devices. To enhance the user experience, the server will retrieve the animal broadcast templates from the immersive data acquisition area corresponding to each data output segment when the user is in an immersive experience. The animal broadcast templates can introduce the living habits of the animals to be viewed, etc. The server will control the voice broadcast unit in the immersive data output area to broadcast the animal broadcast templates. During the immersive experience, users can also interact with the animals based on the interactive unit, such as touching the animal's head. When the server determines that the user has interacted with any animal limb corresponding to the animal located in any data output segment based on the interactive unit, it will generate an interactive action corresponding to the animal limb with a preset number of interactive frames based on the AI ​​module and display it to the user. For example, after the server determines that the user interacts with the head of a tiger based on the interaction unit, it will generate an interactive action corresponding to the head of the tiger based on the AI ​​module with a preset number of interaction frames. For example, the interactive action can be shaking the head and opening the mouth wide. Then, the server will interact with and display this interactive action to the user. Next, the server will retrieve the interactive broadcast template corresponding to the interactive action. For example, the interactive broadcast template can be an explanation of the interactive action. The server will control the voice broadcast unit to broadcast the interactive broadcast template. To avoid extending the user's immersive experience time due to interactive actions, the server will remove each captured image frame located at the end of the data output segment according to the preset number of interactive frames. For example, if the preset number of interactive frames is 5, the server will remove each captured image frame located at the end of the last 5 positions of the data output segment.

[0041] According to the present invention, firstly, the server can accurately determine the first video segment corresponding to the user terminal based on the viewing situation of the user terminal in the immersive data acquisition area at the first location, providing personalized materials for subsequent immersive experiences. Secondly, after the user terminal reaches the immersive data output area at the second location, the server acquires the immersive output points of the first video segment and decomposes them based on the user's points to generate data output segments, enabling subsequent users to view the animals in an immersive way, enhancing the sense of immersion. Simultaneously, the server determines the number of output screens for the corresponding data output segments, which helps to rationally plan screen resources. Finally, when the number of data output screens with idle attributes in the immersive data output area meets the requirements, the server interactively displays each data output segment to the user terminal based on each data output screen, fully utilizing screen resources and providing users with multi-dimensional, interactive display effects, improving user participation and experience.

[0042] Another embodiment of the present invention provides an immersive AI scenic area data output system. Figure 4Its corresponding system block diagram includes: The determination module is configured to determine the first video segment corresponding to each user terminal that meets the immersive data output conditions based on the immersive data acquisition area of ​​each first position. The decomposition module is configured to, in response to any user terminal reaching any corresponding second position of the immersive data output area, acquire the immersive output point of the first video segment and decompose it based on the user point to generate a data output segment, and determine the number of output screens for the corresponding data output segment. The interaction module is configured to respond to a situation where the number of output screens with idle attributes located in the immersive data output area is greater than or equal to the number of output screens, and to interactively display each data output segment to the user terminal based on each of the data output screens.

[0043] In the specification provided herein, the algorithms and displays are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used with the examples of this invention. The required structure for constructing such systems is apparent from the above description. Furthermore, this invention is not directed to any particular programming language. It should be understood that the contents of the invention described herein can be implemented using various programming languages, and the above description of specific languages ​​is for the purpose of disclosing preferred embodiments of the invention.

[0044] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of the invention may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.

[0045] Similarly, it should be understood that, in order to streamline this disclosure and aid in understanding one or more of the various aspects of the invention, in the above description of exemplary embodiments of the invention, various features of the invention are sometimes grouped together in a single embodiment, figure, or description thereof.

[0046] Those skilled in the art will understand that modules, units, or components of the devices disclosed in the examples herein can be arranged in the devices described in this embodiment, or alternatively, can be located in one or more devices different from the devices in this example. The modules in the foregoing examples can be combined into a single module or, in addition, can be divided into multiple sub-modules.

[0047] Those skilled in the art will understand that the modules in the device of the embodiment can be adaptively changed and placed in one or more devices different from that embodiment. Modules, units, or components in the embodiment can be combined into a single module, unit, or component, and further, they can be divided into multiple sub-modules, sub-units, or sub-components.

[0048] Furthermore, those skilled in the art will understand that although some embodiments described herein include certain features included in other embodiments but not others, combinations of features from different embodiments are meant to be within the scope of the invention and form different embodiments.

[0049] Furthermore, some of the embodiments described herein are methods or combinations of method elements that can be implemented by a processor of a computer system or by other means of performing the functions. Therefore, a processor having the necessary instructions for implementing the methods or method elements forms means for implementing the methods or method elements. Furthermore, the elements described herein in the apparatus embodiments are examples of means for implementing the functions performed by elements for the purposes of carrying out the invention.

[0050] As used herein, unless otherwise specified, the use of ordinal numbers such as “first,” “second,” “third,” etc., to describe ordinary objects merely indicates different instances of similar objects and is not intended to imply that the objects being described must have a given order in time, space, ordering, or any other manner.

[0051] Although the invention has been described with respect to a limited number of embodiments, those skilled in the art will understand from the foregoing description that other embodiments are conceivable within the scope of the invention described herein. Furthermore, it should be noted that the language used in this specification has been chosen primarily for readability and edibility purposes, and not for the purpose of explaining or limiting the subject matter of the invention.

Claims

1. An immersive AI-powered scenic area data output method, characterized in that, include: Based on the immersive data acquisition area of ​​each first position, the first video segment corresponding to each user terminal that meets the immersive data output conditions is determined. Upon responding to any user terminal reaching any corresponding second location of the immersive data output area, the immersive output point of the first video segment is acquired and decomposed based on the user point, a data output segment is generated, and the number of output screens for the corresponding data output segment is determined. If the number of output screens with idle attributes located in the immersive data output area is greater than or equal to the number of output screens, the data output segments are interactively displayed to the user terminal based on each of the data output screens. In response to any user terminal arriving at any corresponding second location of the immersive data output area, retrieve each first video segment corresponding to the user terminal; In response to user interaction with any first video segment, retrieve immersive experience images of the immersive data acquisition area corresponding to the first video segment, wherein the immersive experience images include each immersive experience point in the corresponding immersive data acquisition area; In response to user interaction with any immersive experience point based on the immersive experience image, the immersive experience point is determined as the immersive output point. Based on the user point, data output segments corresponding to the user are generated, and the number of output screens corresponding to each data output segment is determined.

2. The method according to claim 1, characterized in that, Based on the immersive data acquisition area of ​​each first location, each first video segment corresponding to each user terminal that meets the immersive data output conditions is determined, including: If the location information of any user terminal overlaps with the location of any immersive data acquisition area, determine the overlapping time period corresponding to the overlapping relationship; If the overlap duration of the corresponding overlapping time period is greater than the preset duration, it is determined that the user terminal meets the immersive data output conditions; Based on the orientation acquisition unit corresponding to the immersive data acquisition area, the orientation acquisition video of the corresponding overlapping time period is obtained, and the orientation acquisition video is recognized based on the video recognition model to obtain the viewing orientation of the corresponding user terminal and the viewing time of each viewing orientation. The viewing orientation corresponding to the maximum viewing time is determined as the user's main orientation at the corresponding user terminal, and the scenic area acquisition unit corresponding to the user's main orientation acquires the scenic area acquisition video of the corresponding overlapping time period. The video footage collected from the scenic area is identified as the first video segment corresponding to the user's device.

3. The method according to claim 1, characterized in that, Based on the user location decomposition, corresponding data output segments are generated for each user terminal, and the number of output screens for each data output segment is determined, including: Based on the acquisition units of each point of the corresponding immersive output point, the acquisition video of each point during the corresponding overlapping time period is obtained, and the target action attribute of each point acquisition video is determined, wherein the target action attribute includes target action and no target action. Based on the preset output time, the video of each point that is captured with the corresponding target attribute of having target action is decomposed into points to obtain each data output segment; The number of data outputs for each corresponding data output segment is determined as the number of output screens.

4. The method according to claim 3, characterized in that, Determine the target action attributes of the video captured at each location, wherein the target action attributes include actions with and without targets, including: Identify the individual image frames that make up the video captured at each location, and perform image recognition on each image frame. Based on the recognition results, the captured image frames of each animal area corresponding to the same location are aggregated into the same animal image group, and the video of the location is updated based on the captured image frames located in the animal image group. If the duration of the video captured at any point after the corresponding video update exceeds a preset time threshold, the target attribute of the video captured at that point is determined to have a target action. Conversely, the target attribute of the video captured at that location is determined to be a targetless action.

5. The method according to claim 4, characterized in that, Based on a preset output time, the video footage captured at each point corresponding to the target attribute of "targeted action" is decomposed into points to obtain each data output segment, including: Retrieve each captured image frame from the video of each point where the target attribute is a target action, and determine the center point of each area of ​​each animal area that indicates each animal to be viewed, located in each captured image frame. Based on the video captured at each location, the center points of each region of the same animal are divided into the same location group according to the images captured in adjacent frames. Each image coordinate system is established with the center point of each acquired image frame as the origin, and the coordinate points of each region corresponding to the center point of each region are determined based on each image coordinate system. The distances between the coordinates of each area corresponding to the same location group are determined based on the distances between the coordinates of each area. If the distance to the response area is greater than the first preset distance, the acquired image frames corresponding to the distance to the area will be aggregated into a moving image group corresponding to the point acquisition video of each acquired image frame. The number of image frames corresponding to the same moving image group is determined based on each acquired image frame located in the same moving image group, and the number of image frames is compared with the number of frames determined based on the preset output time. If the number of response image frames is greater than or equal to a preset number of frames, video fusion is performed on each of the acquired image frames based on the preset number of frames, and point decomposition is performed on each of the resulting data output segments.

6. The method according to claim 5, characterized in that, The method further includes: If the number of response image frames is less than the preset number of frames, determine the number of animal regions in each captured image frame in each captured video frame, excluding the corresponding moving image group. If the number of any region exceeds the preset number, the captured image frame corresponding to that region is added to the moving image group so that the captured image frame is connected to the last captured image frame in the moving image group, thus obtaining a moving image group with updated frame count. If the number of image frames in a moving image group that responds to a frame update is less than a preset number of frames, the acquired image frames whose corresponding moving distance is less than a first preset distance but greater than a second preset distance are added to the moving image group in chronological order until the number of image frames in the corresponding moving image group is greater than or equal to the preset number of frames.

7. The method according to claim 5, characterized in that, The obtained data output segments are decomposed into points, including: Determine the percentage of each image corresponding to each animal region in the output image frames that make up each data output segment, indicating each viewer animal; In response to any image proportion being greater than a first preset proportion and less than a second preset proportion, the animal region corresponding to the image proportion is determined as the first target region, and target extraction is performed on the first target region; The extracted first target region is magnified based on a first magnification factor. Since the magnified first target region overlaps with any animal region, the first target region is reduced in image based on a first reduction factor corresponding to the first magnification factor. If any image proportion is less than or equal to a second preset proportion, the animal region corresponding to the image proportion is determined as the second target region, and target extraction is performed on the second target region; The extracted second target region is magnified based on a second magnification factor, wherein the second magnification factor is greater than the first magnification factor; If the magnified second target region overlaps with any animal region, the second target region is image-reduced based on the corresponding second magnification factor and the second target region is determined as the first target region.

8. The method according to claim 5, characterized in that, Based on the data output screens, each data output segment is interactively displayed to the user terminal, including: Each data output segment is sent to its respective output display slot on each data output screen that has an idle attribute; In response to the user completing the wearing of any interactive unit, the data output segments are output and displayed, and the voice broadcasting unit based on the immersive data output area broadcasts the animal broadcasting templates in the immersive data collection area corresponding to each data output segment. In response to user interaction with any animal limb corresponding to any animal in any data output segment via the interaction unit, the AI ​​module generates an interactive action corresponding to the animal limb with a preset number of interaction frames and displays it to the user. The voice broadcasting unit then broadcasts the interactive broadcasting template corresponding to the interactive action. Based on a preset number of interactive frames, the acquired image frames located at the end of the data output segment are removed.

9. An immersive AI-powered scenic area data output system, characterized in that, include: The determination module is configured to determine the first video segment corresponding to each user terminal that meets the immersive data output conditions based on the immersive data acquisition area of ​​each first position. The decomposition module is configured to, in response to any user terminal reaching any corresponding second position of the immersive data output area, acquire the immersive output point of the first video segment and decompose it based on the user point to generate a data output segment, and determine the number of output screens for the corresponding data output segment. The interaction module is configured to respond to a situation where the number of output screens with idle attributes located in the immersive data output area is greater than or equal to the number of output screens, and to interactively display each data output segment to the user terminal based on each of the data output screens. In response to any user terminal arriving at any corresponding second location of the immersive data output area, retrieve each first video segment corresponding to the user terminal; In response to user interaction with any first video segment, retrieve immersive experience images of the immersive data acquisition area corresponding to the first video segment, wherein the immersive experience images include each immersive experience point in the corresponding immersive data acquisition area; In response to user interaction with any immersive experience point based on the immersive experience image, the immersive experience point is determined as the immersive output point. Based on the user point, data output segments corresponding to the user are generated, and the number of output screens corresponding to each data output segment is determined.