Immersive AI scenic spot data output method and system

An immersive AI-powered approach to scenic area data output, which utilizes personalized video segment generation and dynamic screen resource adjustment, addresses the issues of interactivity and resource utilization in multimedia displays at scenic areas, thereby enhancing visitors' immersive experience and improving resource utilization efficiency.

CN120803271AActive Publication Date: 2025-10-17NANJING MOCHOU INTELLIGENT INFORMATION TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510979282.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-10-17
Estimated Expiration
2045-07-16

AI Technical Summary

Technical Problem

Existing multimedia display methods in scenic areas cannot be customized according to tourists' personalized tour routes and interests, resulting in tourists being unable to effectively interact with the displayed content. Furthermore, the lack of flexibility in screen resource planning leads to idle and wasted resources and an inability to meet peak viewing demand, thus limiting the immersive experience.

Method used

By generating video segments based on user location and interests, and combining this with dynamic adjustments to the immersive data collection area and output screen, a personalized immersive AI scenic area data output method is provided. This includes video recognition, point decomposition, and rational planning of screen resources to achieve interactive display.

Benefits of technology

It enhances the visitor's immersive experience and engagement, makes full use of screen resources, provides multi-dimensional and interactive display effects, and strengthens the user's sense of immersion and experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120803271A_ABST
    Figure CN120803271A_ABST
Patent Text Reader

Abstract

The invention provides an immersive AI scenic spot data output method and system. The method comprises the following steps: determining a first video segment corresponding to each user side meeting an immersion data output condition based on an immersion data acquisition area of each first position; after any user side reaches any immersion data output area corresponding to the second position in response, obtaining an immersion output point location of the first video segment, decomposing the immersion output point location based on the user point location, generating a data output segment, and determining the number of output screens corresponding to the data output segment; and in response to the fact that the number of output screens of the data output screens which are correspondingly located in the immersion data output area and have idle attributes is larger than or equal to the number of the output screens, interactively displaying the data output segments to the user side based on the data output screens. The data output effect is at least improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, in particular to an immersive AI scenic spot data output method and system. BACKGROUND

[0002] With the improvement of people's living standards, tourism has become an important way for the public to relax and entertain. In various scenic spots, tourists no longer satisfy with a superficial tour, but expect to experience the characteristics of the scenic spot in depth with the help of advanced technology, and obtain a unique and interactive play. The demand for immersive tour experience is becoming increasingly strong. Currently, scenic spots generally use traditional multimedia display methods, such as playing standardized promotional videos in a fixed area or setting fixed screens to display scenic landscapes, historical culture and other content. However, such methods cannot be customized according to the personalized tour track and interest preferences of tourists, making it difficult for tourists to interact with the display content effectively. Moreover, the planning of screen resources in scenic spots lacks flexibility, and it is difficult to dynamically adjust the use according to the real-time needs and on-site conditions of tourists, resulting in resource waste and failure to meet the viewing needs of tourists during peak periods, which greatly limits the immersive experience of tourists. SUMMARY

[0003] Based on the above problems, the present application is proposed to provide an immersive AI scenic spot data output method and system to overcome the above problems or at least partially solve the above problems.

[0004] According to one aspect of the present application, an immersive AI scenic spot data output method is provided, comprising the following steps: determining, based on the immersive data collection area of each first position, a first video segment corresponding to each user terminal that meets the immersive data output condition; in response to any of the user terminals reaching the immersive data output area of any corresponding second position, obtaining the first video segment immersive output point and decomposing based on the user point to generate a data output segment, and determining the number of output screens corresponding to the data output segment; in response to the number of output screens of each data output screen corresponding to the immersive data output area having an idle attribute being greater than or equal to the number of output screens, each data output segment is interactively displayed to the user terminal based on each data output screen.

[0005] Optionally, in the method according to the present application, determining, based on the immersive data collection area of each first position, each first video segment corresponding to each user terminal that meets the immersive data output condition comprises: in response to the positional information of any user terminal having an overlapping relationship with the area position of any immersive data collection area, determining an overlapping period corresponding to the overlapping relationship; in response to the overlap duration corresponding to the overlap period being greater than a preset duration, determining that the user terminal satisfies the immersive data output condition; acquiring, based on the orientation acquisition unit corresponding to the immersive data acquisition area, an orientation acquisition video corresponding to the overlap period, and performing video recognition on the orientation acquisition video based on a video recognition model to obtain each viewing orientation corresponding to the user terminal and each viewing time corresponding to each viewing orientation; determining the viewing orientation corresponding to the maximum viewing time as the user main orientation corresponding to the user terminal, and acquiring, based on the scenic area acquisition unit corresponding to the user main orientation, a scenic area acquisition video corresponding to the overlap period; determining the scenic area acquisition video as the first video segment corresponding to the user terminal.

[0006] Optionally, in the method according to the present application, in response to any of the user terminals reaching the immersive data output area of any corresponding second position, the first video segment immersive output point is acquired and decomposed based on the user point to generate each data output segment, and the output screen number corresponding to each data output segment is determined, comprising: in response to any of the user terminals reaching the immersive data output area of any corresponding second position, retrieving each first video segment corresponding to the user terminal; in response to the user terminal interacting with any first video segment, retrieving an immersive experience image of the immersive data acquisition area corresponding to the first video segment, wherein the immersive experience image comprises each immersive experience point of the immersive data acquisition area; in response to the user terminal interacting with any immersive experience point based on the immersive experience image, determining the immersive experience point as the immersive output point, generating each data output segment corresponding to the user terminal based on the user point decomposition, and determining the output screen number corresponding to each data output segment.

[0007] Optionally, in the method according to the present application, generating each data output segment corresponding to the user terminal based on the user point decomposition and determining the output screen number corresponding to each data output segment comprises: acquiring each point acquisition video corresponding to the overlap period based on each point acquisition unit corresponding to the immersive output point, and determining the target action attribute of each point acquisition video, wherein the target action attribute comprises a target action and no target action; performing point decomposition on each point acquisition video corresponding to the target attribute of the target action based on a preset output time to obtain each data output segment; determining the data output number corresponding to each data output segment as the output screen number.

[0008] Optionally, in the method according to the present application, determining the target action attribute of each point acquisition video, wherein the target action attribute comprises a target action and no target action, comprises: determining each acquisition image frame constituting each point position acquisition video, and performing image recognition on each acquisition image frame; based on the recognition result, each acquisition image frame indicating each animal region of any ornamental animal existing in the same point position acquisition video is summarized to the same animal image group, and each acquisition image frame located in the animal image group is used to update the point position acquisition video; in response to the video time length of the updated video of any point position being greater than a preset time threshold, the target attribute of the point position acquisition video is determined as having a target action; otherwise, the target attribute of the point position acquisition video is determined as not having a target action.

[0009] Optionally, in the method according to the present application, each point position acquisition video corresponding to the target attribute of having a target action is decomposed based on a preset output time to obtain each data output segment, including: each acquisition image frame corresponding to each point position acquisition video with the target attribute of having a target action is called, and each region center point indicating each animal region of each ornamental animal located in each acquisition image frame is determined; based on each point position acquisition video, each region center point indicating the same ornamental animal of each acquisition image frame with adjacent relationship is divided into the same point position group; each image coordinate system is established with each image center point of each acquisition image frame as the origin, and each region coordinate point corresponding to each region center point is determined based on each image coordinate system; based on each coordinate distance of each region coordinate point corresponding to the same point position group, each region distance corresponding to each animal region of the same ornamental animal is determined; in response to the region distance being greater than a first preset distance, each acquisition image frame corresponding to the region distance is summarized to a moving image group corresponding to the point position acquisition video corresponding to each acquisition image frame; based on each acquisition image frame located in the same moving image group, the image frame number corresponding to the moving image group is determined, and the image frame number is compared with the preset frame number determined based on the preset output time; in response to the image frame number being greater than or equal to the preset frame number, each acquisition image frame is video fused based on the preset frame number, and each data output segment obtained is decomposed by point position.

[0010] Optionally, in the method according to the present application, the method further comprises: in response to the image frame number being less than the preset frame number, the region number of each animal region located in each acquisition image frame except the acquisition image frame corresponding to the moving image group in each point position acquisition video is determined; In response to any of the region numbers being greater than a preset number, the collected image frames corresponding to the region numbers are added to the moving image group so that the collected image frames are connected to the last collected image frame in the moving image group, and a moving image group with updated frame numbers is obtained; In response to the number of image frames of the moving image group with updated frame numbers being less than a preset number of frames, each collected image frame corresponding to a moving distance less than a first preset distance and greater than a second preset distance is sequentially added to the moving image group in chronological order until the number of image frames of the moving image group is greater than or equal to the preset number of frames.

[0011] Optionally, in the method according to the present application, the obtained data output segments are subjected to point position decomposition, comprising: determining image proportions corresponding to each animal region indicating each viewing animal in the output image frames constituting each data output segment; In response to any of the image proportions being greater than a first preset proportion and less than a second preset proportion, the animal region corresponding to the image proportion is determined as a first target region, and target extraction is performed on the first target region; The extracted first target region is subjected to image enlargement based on a first enlargement coefficient, and in response to the enlarged first target region having an overlapping relationship with any animal region, the first target region is subjected to image reduction based on a first reduction coefficient corresponding to the first enlargement coefficient; In response to any of the image proportions being less than or equal to the second preset proportion, the animal region corresponding to the image proportion is determined as a second target region, and target extraction is performed on the second target region; The extracted second target region is subjected to image enlargement based on a second enlargement coefficient, wherein the second enlargement coefficient is greater than the first enlargement coefficient; In response to the enlarged second target region having an overlapping relationship with any animal region, the second target region is subjected to image reduction based on a second reduction coefficient corresponding to the second enlargement coefficient, and the second target region is determined as the first target region.

[0012] Optionally, in the method according to the present application, each data output segment is interactively displayed to the user end based on each data output screen, comprising: Each data output segment is sent to each output display slot of each data output screen having an idle attribute; In response to the user completing the wearing of any interactive unit, each data output segment is output displayed, and an animal broadcast template of an immersive data collection region corresponding to each data output segment is subjected to voice broadcast based on a voice broadcast unit of an immersive data output region; In response to the user interacting with any animal limb corresponding to the ornamental animal located at any data output section based on the interaction unit, the AI module generates an interaction action corresponding to the animal limb with a preset number of interaction frames and interacts with the user, and the voice broadcast unit performs voice broadcast on an interaction broadcast template corresponding to the interaction action; Based on the preset number of interaction frames, the removal operation is performed on each collected image frame located at the end of the data output section.

[0013] According to another aspect of the present application, an immersive AI scenic spot data output system is provided, comprising: The determination module is configured to determine a first video section corresponding to each user terminal satisfying the immersive data output condition based on the immersive data collection area of each first position; The decomposition module is configured to, in response to any user terminal reaching the immersive data output area of any corresponding second position, acquire the immersive output point of the first video section and decompose based on the user point to generate a data output section, and determine the output screen number of the corresponding data output section; The interaction module is configured to, in response to the output screen number of each data output screen corresponding to the immersive data output area with an idle attribute being greater than or equal to the output screen number, interact with each data output section to the user terminal based on each data output screen.

[0014] According to the scheme of the present application, first, the server can accurately determine the first video section corresponding to the user terminal according to the observation of the user terminal in the immersive data collection area of the first position satisfying the immersive data output condition, providing personalized materials for subsequent immersive experience. Second, after the user terminal reaches the immersive data output area of the second position, the server will acquire the immersive output point of the first video section and decompose based on the user point to generate a data output section, so that subsequent users can observe the ornamental animal in situ, enhancing the sense of immersion. At the same time, the server will determine the output screen number of the corresponding data output section, which is helpful for reasonable planning of screen resources. Finally, when the number of data output screens with idle attributes in the immersive data output area meets the requirements, the server will interact with each data output section to the user terminal according to each data output screen, which not only makes full use of screen resources, but also provides multi-dimensional and interactive display effects for users, improving user participation and experience. BRIEF DESCRIPTION OF DRAWINGS

[0015] Figure 1 A flowchart of an immersive AI scenic spot data output method according to an embodiment of the present application is shown; Figure 2 A schematic diagram of an immersive experience point according to an embodiment of the present application is shown; Figure 3 A schematic diagram of a region distance is shown according to one embodiment of the present application; Figure 4 A structural block diagram of an immersive AI scenic spot data output system according to another embodiment of the present application is shown. DETAILED DESCRIPTION

[0016] Exemplary embodiments of the present disclosure will be described in greater detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be accurately conveyed to those skilled in the art.

[0017] To solve the problems in the background art, the inventors propose the solutions of the present application. One embodiment of the present application provides an immersive AI scenic spot data output method, which can be executed in a computing device.

[0018] Figure 1 A flowchart of an immersive AI scenic spot data output method according to one embodiment of the present application is shown, which is suitable for execution in a computing device.

[0019] As Figure 1 shown, the immersive AI scenic spot data output method proposed in this embodiment starts from step S102, in which the following contents are included: determining, based on the immersive data collection region at each first position, a first video segment corresponding to each user terminal that meets the immersive data output condition.

[0020] For example, in this embodiment, the user terminal can be understood as a tourist wearing a positioning bracelet, and there are multiple immersive data collection regions at each first position in the scenic spot, which can be understood as different animal pavilions in the zoo; In each immersive data collection region in this embodiment, not only are there different animals to watch, but also the first video segment corresponding to the user terminal can be determined according to the viewing situation of the user terminal in the immersive data collection region that meets the immersive data output condition. The generation of the first video segment can facilitate the subsequent server to generate each data output segment for the user to experience immersion.

[0021] Further, the above-mentioned "determining, based on the immersive data collection region at each first position, a first video segment corresponding to each user terminal that meets the immersive data output condition" further includes the following steps: in response to the positional information of any user terminal and the region position of any immersive data collection region having an overlapping relationship, determining an overlapping time period corresponding to the overlapping relationship; In response to the overlap duration of the corresponding overlap period being greater than a preset duration, it is determined that the user terminal satisfies the immersive data output condition; Based on the orientation acquisition unit corresponding to the immersive data acquisition area, an orientation acquisition video corresponding to the overlap period is acquired, and a video recognition model is used to perform video recognition on the orientation acquisition video to obtain each viewing orientation corresponding to the user terminal and each viewing time corresponding to each viewing orientation; The viewing orientation corresponding to the maximum viewing time is determined as the user main orientation corresponding to the user terminal, and the scenic area acquisition unit corresponding to the user main orientation is used to acquire a scenic area acquisition video corresponding to the overlap period; The scenic area acquisition video is determined as the first video segment corresponding to the user terminal.

[0022] For example, in this embodiment, the server acquires the position information of the user terminal in real time, and when the position information of any one user terminal has an overlap relationship with the area position of any one immersive data acquisition area, it indicates that the user terminal has entered the immersive data acquisition area; Next, the server determines the duration that the user terminal is in the immersive data acquisition area, that is, the overlap duration of the overlap period when the position information has an overlap relationship with the area position, and when the overlap duration is greater than a preset duration, it indicates that the user terminal is in the immersive data acquisition area for a long time, that is, it can be preliminarily determined that the user terminal is more interested in the viewing animals located in the immersive data acquisition area. Therefore, the server determines that the user terminal satisfies the immersive data output condition; In order to make the subsequent generated data output segment more in line with the viewing situation of the user terminal, the server will first determine which direction the user terminal is more interested in when viewing animals by the following method; First, the server acquires an orientation acquisition video corresponding to the overlap period through the orientation acquisition unit arranged in the immersive data acquisition area, and then performs video recognition on the orientation acquisition video through a video recognition model, so as to obtain each viewing orientation corresponding to the user terminal and each viewing time corresponding to each viewing orientation corresponding to the overlap period; Next, the server determines the viewing orientation corresponding to the maximum viewing time as the user main orientation corresponding to the user terminal, for example, the obtained each viewing orientation corresponding to the user terminal is east, west, and the corresponding each viewing time is 10 minutes and 15 minutes respectively. The server determines the west viewing orientation as the user main orientation corresponding to the user terminal; Finally, the server acquires a scenic area acquisition video corresponding to the overlap period according to the scenic area acquisition unit corresponding to the user main orientation, and determines the scenic area acquisition video as the first video segment corresponding to the user terminal, so that the user terminal can quickly select the experience scene that wants to experience immersive experience based on each first video segment.

[0023] In step S104, the following is included: In response to any of the user terminals reaching any corresponding second position of the immersive data output area, the immersive output point of the first video segment is obtained and decomposed based on the user point to generate a data output segment, and the number of output screens corresponding to the data output segment is determined.

[0024] For example, in this embodiment, multiple immersive data output areas will be set up in the scenic spot, i.e. areas for subsequent user terminals to experience immersion. When a user terminal reaches any corresponding second position of the immersive data output area, the server will obtain the immersive output point of the first video segment and decompose it based on the user point to generate a data output segment. Since the immersive output point is set close to the animal for viewing, subsequent user terminals can experience the animal in situ according to the data output segment. There can be multiple data output screens in the immersive data output area, and the number of output screens corresponding to different user terminals is different. Therefore, the server determines the number of output screens corresponding to the data output segment to facilitate subsequent screen planning for each data output screen, so that each user terminal can quickly and efficiently experience immersion.

[0025] Further, the above-mentioned "in response to any of the user terminals reaching any corresponding second position of the immersive data output area, obtaining the immersive output point of the first video segment and decomposing it based on the user point to generate each data output segment, and determining the number of output screens corresponding to each data output segment" further includes the following steps: In response to any of the user terminals reaching any corresponding second position of the immersive data output area, each first video segment corresponding to the user terminal is retrieved. In response to the user terminal interacting with any first video segment, the immersive experience image of the immersive data collection area corresponding to the first video segment is retrieved, wherein the immersive experience image includes each immersive experience point of the immersive data collection area. In response to the user terminal interacting with any immersive experience point based on the immersive experience image, the immersive experience point is determined as the immersive output point, each data output segment corresponding to the user terminal is generated based on the user point decomposition, and the number of output screens corresponding to each data output segment is determined.

[0026] For example, in this embodiment, when any user terminal reaches any corresponding second position of the immersive data output area, it can be understood that the user terminal wants to experience immersion. At this time, the server will retrieve each first video segment corresponding to the user terminal, so that the user terminal can select the specific scene it wants to experience immersion by viewing each first video segment. When the user terminal interacts with any one of the first video segments, the server will call the immersive experience image of the immersive data collection area corresponding to the first video segment. The immersive experience image can be understood as a top view of the area corresponding to the animal area of the ornamental animal located in the immersive data collection area, and there are various immersive experience points in the immersive experience image, such as Figure 2 When the user terminal interacts with any one of the immersive experience points in the immersive experience image, it indicates that the user terminal wants to immerse in the point of view of the immersive experience point, so the server will determine the immersive experience point as the immersive output point, and then perform user point decomposition to generate various data output segments corresponding to the user terminal, and determine the output screen quantity corresponding to each data output segment, so as to facilitate subsequent reasonable arrangement for the immersive experience of the user terminal.

[0027] Further, the above-mentioned "generating various data output segments corresponding to the user terminal based on user point decomposition, and determining the output screen quantity corresponding to each data output segment" further includes the following steps: Based on the various point collection units corresponding to the immersive output point, the various point collection videos corresponding to the overlapping time period are obtained, and the target action attribute of each point collection video is determined, wherein the target action attribute includes target action and no target action; Based on the preset output time, the point decomposition is performed on the various point collection videos corresponding to the target attribute of the target action, and various data output segments are obtained; The data output quantity corresponding to each data output segment is determined as the output screen quantity.

[0028] For example, in the present embodiment, in order to enable the data output segments generated subsequently for the immersive output point to give the user terminal a more realistic and rich immersive experience, a plurality of collection units are arranged corresponding to one immersive output point, so as to facilitate the omnidirectional video collection of the ornamental animals located in the immersive collection area based on the immersive output point; In order to ensure that the data output segments generated subsequently are consistent with the video scene of the first video segment of the user terminal, the server will obtain the various point collection videos corresponding to the overlapping time period based on the various point collection units corresponding to the immersive output point; If there is no ornamental animal or the ornamental animal exists only temporarily in the data output segments generated subsequently, when the user terminal performs immersive experience, the corresponding immersive experience will be greatly reduced, and unnecessary screen consumption will also be increased, so the server will first determine the target action attribute of each point collection video, specifically, the target action attribute includes target action and no target action; ​In order to improve the experience efficiency of each user terminal in immersive experience, the server will collect video points of each point position corresponding to the target attribute of the target action according to the preset output time, so as to obtain each data output segment, and the preset output time can be set in advance by the management terminal according to the actual situation. Finally, the server will determine the data output quantity corresponding to each data output segment as the output screen quantity, so as to facilitate subsequent planning of each data output screen, so that the user terminal can quickly have an immersive experience.

[0029] Furthermore, the above "determining the target action attribute of each point position collected video, wherein the target action attribute includes target action and non-target action" further includes the following steps: determining each collected image frame constituting each point position collected video, and performing image recognition on each collected image frame; Based on the recognition result, each collected image frame of each animal region indicating any ornamental animal corresponding to the same point position collected video is summarized to the same animal image group, and each collected image frame located in the animal image group is used to update the point position collected video; In response to the video time length of the video updated at any point position collected video being greater than a preset time threshold, the target attribute of the point position collected video is determined as target action; On the contrary, the target attribute of the point position collected video is determined as non-target action.

[0030] For example, in the present embodiment, the server will first determine each collected image frame constituting each point position collected video, and then perform image recognition on each collected image frame to determine whether there is an animal region indicating an ornamental animal in each collected image frame; Then, the server will summarize each collected image frame of each animal region indicating any ornamental animal corresponding to the same point position collected video to the same animal image group according to the recognition result, and update the point position collected video according to each collected image frame located in the animal image group, so that there is an animal region indicating an ornamental animal in each collected image frame constituting the updated point position collected video; If the video time length of the updated point position collected video is too short, that is, it cannot provide the user terminal with an immersive experience of sufficient time length, the experience of the user terminal will be reduced, so the server will compare the video time length of the point position collected video with a preset time threshold; One possible comparison result is that the video time length is greater than the preset time threshold, which indicates that the video time length of the point position collected video is relatively long, that is, the user terminal can be provided with an immersive experience of sufficient time length based on the point position collected video in the subsequent process, so the server will determine the target attribute of the point position collected video as target action; Another possible comparison result is that the video duration is less than or equal to a preset time threshold. This comparison result indicates that the video duration of the corresponding point collection video is short, that is, subsequent immersive experience of sufficient duration cannot be provided to the user end based on the point collection video, and therefore the server determines the target attribute of the point collection video as no target action.

[0031] Furthermore, the above "point site decomposition of each point collection video with a corresponding target attribute of target action based on a preset output time, to obtain each data output segment" further includes the following steps: Each image frame of each point collection video with a corresponding target attribute of target action is called, and each region center point indicating each animal region of each ornamental animal located in each image frame is determined; Each region center point indicating the same ornamental animal of each image frame having an adjacent relationship is divided into the same point site group based on each point collection video; Each image center point of each image frame is taken as the origin to establish each image coordinate system, and each region coordinate point corresponding to each region center point is determined based on each image coordinate system; Each region distance corresponding to each animal region indicating the same ornamental animal is determined based on each coordinate distance of each region coordinate point corresponding to the same point site group; In response to the region distance being greater than a first preset distance, each image frame corresponding to the region distance is summarized to a moving image group corresponding to the point collection video corresponding to each image frame; The image frame number corresponding to the same moving image group is determined based on each image frame located in the same moving image group, and frame number comparison is performed between the image frame number and a preset frame number determined based on a preset output time; In response to the image frame number being greater than or equal to the preset frame number, video fusion is performed on each image frame based on the preset frame number, and point site decomposition is performed on each data output segment obtained.

[0032] For example, in the present embodiment, it can be understood that when the user is immersed in the experience, it is more desirable to see the ornamental animals in the moving process rather than the ornamental animals in the stationary state, and therefore the server will first determine the moving speed of the ornamental animals located in each point collection video; First, the server will first call each image frame of each point collection video with a corresponding target attribute of target action, and then determine each region center point indicating each animal region of each ornamental animal located in each image frame; Next, the server will group the center points of the regions of the adjacent captured image frames indicating the same animal into the same point group according to the captured videos at each point. It will then establish image coordinate systems with the center points of the images of the captured image frames as the origins, and determine the coordinate points of the regions corresponding to the center points of the regions based on the image coordinate systems. Next, the server determines the coordinate distance between the two regional coordinate points according to the regional coordinate points of the two regional center points corresponding to the same point group, and determines each coordinate distance as the regional distance of each animal area corresponding to the same viewing animal, such as Figure 3 As shown; The larger the regional distance, the faster the movement speed of the corresponding ornamental animal. Therefore, the server aggregates each captured image frame whose corresponding regional distance is greater than the first preset distance into a moving image group corresponding to the point capture video corresponding to each captured image frame. Then, the server determines the image frame number of each captured image frame in the same moving image group, and then determines the preset frame number corresponding to the preset output time. At this time, the server compares the image frame number with the preset frame number; When the number of image frames is greater than or equal to the preset number of frames, it means that the video duration of the video composed of the corresponding number of image frames can meet the subsequent immersive experience requirement of the user end. Therefore, the server will perform video fusion on each captured image frame according to the preset number of frames to obtain a data output segment corresponding to the preset output time; Finally, in order to improve the immersive experience of subsequent users, the server will further decompose each data output segment obtained.

[0033] Furthermore, the above method further includes the following steps: In response to the number of image frames being less than a preset number of frames, determining the number of regions of each animal region in each captured image frame excluding the corresponding moving image group in the captured video at each point; In response to the presence of any region having a number greater than a preset number, adding captured image frames corresponding to the number of regions to the moving image group so that the captured image frames are connected to the last captured image frame in the moving image group, thereby obtaining a moving image group with an updated number of frames; In response to the image frame number of the moving image group corresponding to the frame number update being less than the preset frame number, each captured image frame whose corresponding moving distance is less than the first preset distance and greater than the second preset distance is added to the moving image group in sequence based on chronological order until the image frame number of the corresponding moving image group is greater than or equal to the preset frame number.

[0034] For example, in the embodiment, when the image frame number is less than the preset frame number, it indicates that the video time length of the video composed of the corresponding image frames cannot reach the time length requirement of subsequent immersive experience of the user end, at this time, the server will further perform image screening on the remaining each collection image frame, and add the collection image frame meeting the requirement to the moving image group, so that the image frame number of the corresponding moving image group is greater than or equal to the preset frame number; In order to enable the subsequent user end to see more ornamental animals during the immersive experience, the server will determine the region number of each animal region of each collection image frame in the collection video at each point in addition to the corresponding moving image group; When the region number is greater than the preset number, it indicates that the region number of the animal region corresponding to the collection image frame is relatively large, so the server will add the collection image frame corresponding to the region number to the moving image group, and add the collection image frame to the last collection image frame in the moving image group, to obtain the moving image group with updated frame number, which can maximize avoid removing the collection image frame of the ornamental animal in the moving state when the subsequent removal operation is performed on each collection image frame at the end according to the preset interaction frame number based on the data output segment determined based on the moving image group; If the image frame number of the moving image group with updated frame number is still less than the preset frame number, the server will further add each collection image frame corresponding to a moving distance less than the first preset distance and greater than the second preset distance to the moving image group in time sequence until the image frame number of the moving image group is greater than or equal to the preset frame number.

[0035] Furthermore, the above-mentioned "point position decomposition of the obtained each data output segment" further includes the following steps: Determine each image proportion corresponding to each animal region indicating each ornamental animal in the output image frame constituting each data output segment; In response to any image proportion being greater than a first preset proportion and less than a second preset proportion, determine the animal region corresponding to the image proportion as a first target region, and perform target extraction on the first target region; Perform image magnification on the extracted first target region based on a first magnification coefficient, and in response to the magnified first target region having an overlapping relationship with any animal region, perform image reduction on the first target region based on a first reduction coefficient corresponding to the first magnification coefficient; In response to any image proportion being less than or equal to the second preset proportion, determine the animal region corresponding to the image proportion as a second target region, and perform target extraction on the second target region; Perform image magnification on the extracted second target region based on a second magnification coefficient, wherein the second magnification coefficient is greater than the first magnification coefficient; In response to the second target region after amplification having an overlapping relationship with any animal region, the second target region is image-reduced based on a second reduction coefficient corresponding to a second amplification coefficient, and the second target region is determined as the first target region.

[0036] For example, in the present embodiment, since the distance between the ornamental animals and the collection unit can be different, it leads to that the image proportion of each animal region indicating any ornamental animal in each output image frame constituting each data output segment and the output image frame corresponding thereto is different; When the image proportion is too small, it will lead to that the subsequent user end cannot clearly see the ornamental animal when performing immersive experience, in order to avoid reducing the experience of the user end, the server will first determine each image proportion corresponding to each animal region located in the output image frame constituting each data output segment; In order to make the subsequent image proportion corresponding to each animal region be relatively close, facilitate the consistency of the user end to each ornamental animal for immersive experience, the management end can set a first preset proportion and a second preset proportion in the server in advance; One possible condition is that the image proportion is greater than the first preset proportion and less than the second preset proportion, that is, the animal region corresponding to the image proportion is not very small, only needs to be amplified by a small amplitude, at this time the server will determine the animal region corresponding to the image proportion as the first target region, and perform target extraction on the first target region; Then, the server will image amplify the extracted first target region according to the first amplification coefficient, in order to ensure that the subsequent user end can see the complete animal region when performing immersive experience, the server will further determine whether the amplified first target region has an overlapping relationship with other animal regions, and when having the overlapping relationship, the first target region is image-reduced based on a first reduction coefficient corresponding to the first amplification coefficient, that is, the image proportion of the reduced first target region is the initial image proportion; Another possible condition is that the image proportion is less than or equal to the second preset proportion, that is, the animal region corresponding to the image proportion is too small, which needs to be amplified by a large amplitude, therefore the server will determine the animal region corresponding to the image proportion as the second target region, and then perform target extraction on the second target region; Then, the server will image amplify the extracted second target region according to the second amplification coefficient, it can be explained that the second amplification coefficient is greater than the first amplification coefficient; Similarly, the server will further determine whether the amplified second target region has an overlapping relationship with other animal regions, and when having the overlapping relationship, the second target region is image-reduced based on a second reduction coefficient corresponding to the second amplification coefficient; In order to maximize the image zooming of the second target area, the server determines the second target area as the first target area, and performs image zooming on the second target area based on the first zooming coefficient in the above manner.

[0037] In step S106, the following is included: In response to the number of output screens of each data output screen with an idle attribute located in the immersion data output area being greater than or equal to the output screen number, each data output segment is interactively displayed to the user terminal based on each data output screen.

[0038] For example, in the present embodiment, there can be multiple data output screens in the immersion data output area. When the number of output screens of each data output screen with an idle attribute is greater than or equal to the output screen number, it indicates that the immersion data output area has already met the screen number condition for providing an immersive experience to the user terminal. At this time, the server will interactively display each data output segment to the user terminal based on each data output screen.

[0039] Further, the above-mentioned "interactively displaying each data output segment to the user terminal based on each data output screen" further includes the following steps: Each data output segment is sent to each output display slot of each data output screen with an idle attribute; In response to the user completing the wearing of any interactive unit, each data output segment is output displayed, and the voice broadcast template of the animal broadcast module corresponding to the immersion data collection area of each data output segment is voice broadcasted based on the voice broadcast unit of the immersion data output area; In response to the user interacting with any animal limb corresponding to the ornamental animal located in any data output segment based on the interactive unit, the AI module generates an interactive action corresponding to the animal limb with a preset number of interaction frames and interactively displays it to the user, and the voice broadcast unit voice broadcasts the interaction broadcast template corresponding to the interactive action; Each collection image frame located at the end of the data output segment is removed based on the preset number of interaction frames.

[0040] For example, in the present embodiment, the server will send each data output segment to each output display slot of each data output screen with an idle attribute, and after the user completes the wearing of any interactive unit, each data output segment is output displayed to the user, i.e., the user starts the immersive interaction. Here, the interactive unit can be understood as an interactive device such as a VR glasses, a VR glove, etc. In order to improve the experience of a certain user, the server will call the animal broadcast template of the immersion data collection area corresponding to each data output section when the user is immersed, and the animal broadcast template can be used to introduce the living habits of the ornamental animals, etc. The server will control the voice broadcast unit of the immersion data output area to perform voice broadcast on the animal broadcast template. The user can also interact with the ornamental animals based on the interaction unit during the immersion experience, such as touching the head of the ornamental animals. When the server determines that the user interacts with any animal limb corresponding to the ornamental animal located in any one data output section based on the interaction unit, the server will generate an interaction action corresponding to the animal limb based on the AI module and interact with the user for a preset number of interaction frames. For example, after the server determines that the user interacts with the head of the ornamental animal, which is a tiger, based on the interaction unit, the server will generate an interaction action corresponding to the head based on the AI module, such as shaking the head and opening the mouth. Then, the server will interact with the user by displaying the interaction action. Then, the server will call the interaction broadcast template corresponding to the interaction action, such as a commentary on the interaction action, etc. The server will control the voice broadcast unit to perform voice broadcast on the interaction broadcast template. In order to avoid prolonging the immersion experience time of the user due to the interaction action, the server will remove each collection image frame located at the last position of the data output section according to a preset number of interaction frames. For example, if the preset number of interaction frames is 5, the server will remove each collection image frame located at the last 5 positions of the data output section.

[0041] According to the scheme of the present application, first, the server can accurately determine the first video section corresponding to the user end according to the ornamental situation of the immersion data collection area at the first position of the user end that meets the immersion data output condition, and provide personalized materials for subsequent immersion experience. Second, after the user end reaches the immersion data output area at the second position, the server will obtain the first video section immersion output point and decompose based on the user point to generate a data output section, so that the user can observe the ornamental animals in person, and enhance the immersion experience. At the same time, the server will determine the output screen number corresponding to the data output section, which is helpful for reasonable planning of screen resources. Finally, when the number of idle data output screens of the immersion data output area meets the requirements, the server will interact with the user end according to each data output screen, which not only makes full use of the screen resources, but also provides multi-dimensional and interactive display effect for the user, and improves the participation and experience of the user.

[0042] Another embodiment of the present application provides an immersive AI scenic area data output system, Figure 4For its corresponding system block diagram, the system comprises: A determination module configured to determine, based on the immersion data collection area of each first position, a first video segment corresponding to each user terminal satisfying the immersion data output condition A decomposition module configured to, in response to any of the user terminals reaching the immersion data output area of any corresponding second position, acquire the first video segment immersion output point and decompose based on the user point to generate a data output segment, and determine the output screen number of the corresponding data output segment; An interaction module configured to, in response to the output screen number of each data output screen with an idle attribute corresponding to the immersion data output area being greater than or equal to the output screen number, interactively display each data output segment to the user terminal based on each data output screen.

[0043] In the specification provided herein, the algorithms and displays are not inherently related to any particular computer, virtual system, or other apparatus. Various general purpose systems can be used with examples of the present invention. Structured as required by the description above, the structure required to construct such a system is apparent from the above description. Furthermore, the present invention is not directed to any particular programming language. It will be appreciated that a variety of programming languages can be used to implement the present invention described herein, and that the descriptions above are provided for the best mode for carrying out the present invention.

[0044] In the specification provided herein, a large number of specific details are described. However, it can be understood that embodiments of the present invention can be practiced without these specific details. In some instances, well-known methods, structures and techniques have not been described in detail in order not to obscure the understanding of this specification.

[0045] Similarly, it is to be understood that the various features of the inventive aspects sometimes described in the specification in the description of one or more of the example embodiments of the invention are, of course, combinable. There can be many alterations to the embodiments of this invention described herein without departing from the essential scope of the application. Accordingly, the scope of the present invention is to be limited only by the appended claims.

[0046] Those skilled in the art will understand that the modules, or units, or components of the devices in the examples disclosed herein can be arranged in the devices as described in the examples, or alternatively can be located in one or more devices different from the devices in the examples. The modules in the foregoing examples can be combined into one module or further divided into multiple sub-modules.

[0047] Those skilled in the art can understand that the modules in the devices in the examples can be adaptively changed and arranged in one or more devices different from the examples. The modules or units or components in the examples can be combined into one module or unit or component, and further divided into multiple sub-modules or sub-units or sub-components.

[0048] Furthermore, to the extent that the terms "including", "includes", "having", "has", "with", or variants thereof are used in either the detailed description and / or the claims section, such terms are typically intended to be inclusive in a manner similar to the term "comprising" as an open transition term without precluding any additional or

[0049] Furthermore, some of the embodiments described herein are of a "method" or a "process" that can be embodied in software, firmware or hardware, and / or a combination of software, firmware or hardware. Furthermore, some of the embodiments described herein are of a "device" or an "apparatus" that can be implemented in software, firmware or hardware, and / or a combination of software, firmware or hardware. Therefore, "method" and "process" embodiments can be implemented in software, firmware or hardware, and / or a combination of software, firmware or hardware. Similarly, "device" and "apparatus" embodiments can be implemented in software, firmware or hardware, and / or a combination of software, firmware or hardware.

[0050] As used herein, unless otherwise indicated, the use of the ordinal adjectives "first", "second", "third", etc., merely to distinguish different instances of an object to which the adjective refers, and are not intended to denote a

[0051] Although the application has been described and illustrated with respect to a limited number of embodiments, those skilled in the art will appreciate that various modifications can be made without departing from the scope of the application. Accordingly, the scope of the application is limited only by the following claims.

Claims

1. An immersive AI scenic spot data output method, characterized in that: include: Determining, based on the immersive data collection area of ​​each first position, a first video segment corresponding to each user terminal that meets the immersive data output condition; In response to any of the user terminals arriving at any immersive data output area corresponding to the second position, acquiring the immersive output point of the first video segment and decomposing it based on the user point to generate a data output segment, and determining the number of output screens corresponding to the data output segment; In response to the number of output screens corresponding to the data output screens with an idle attribute located in the immersive data output area being greater than or equal to the number of output screens, the data output segments are interactively displayed to the user terminal based on the data output screens.

2. The method according to claim 1, characterized in that Determining first video segments corresponding to user terminals that meet immersive data output conditions based on immersive data collection areas at first locations includes: In response to the position information corresponding to any user terminal and the area position corresponding to any immersive data collection area having an overlapping relationship, determining an overlapping period corresponding to the overlapping relationship; In response to the overlapping duration of the corresponding overlapping period being greater than a preset duration, determining that the user terminal meets the immersive data output condition; Acquiring orientation collection videos corresponding to the overlapping time periods based on the orientation collection unit corresponding to the immersive data collection area, and performing video recognition on the orientation collection videos based on the video recognition model to obtain viewing orientations of the corresponding user terminal and viewing times corresponding to the viewing orientations; Determine the viewing direction corresponding to the maximum viewing time as the user main direction of the corresponding user terminal, and obtain the scenic spot collection video of the corresponding overlapping period based on the scenic spot collection unit corresponding to the user main direction; The scenic spot captured video is determined as a first video segment corresponding to the user terminal.

3. The method according to claim 2, characterized in that In response to any of the user terminals arriving at any immersive data output area corresponding to the second position, obtaining the immersive output point of the first video segment and decomposing it based on the user point to generate each data output segment, and determining the number of output screens corresponding to each data output segment, including: In response to any of the user terminals arriving at any immersive data output area corresponding to the second position, retrieving each first video segment corresponding to the user terminal; In response to a user terminal's interaction with any first video segment, retrieving an immersive experience image of the immersive data collection area corresponding to the first video segment, wherein the immersive experience image includes each immersive experience point corresponding to the immersive data collection area; In response to the user terminal's interaction with any immersive experience point based on the immersive experience image, the immersive experience point is determined as an immersive output point, data output segments corresponding to the user terminal are generated based on the user point decomposition, and the number of output screens corresponding to each data output segment is determined.

4. The method according to claim 3, characterized in that Generate data output segments corresponding to the user end based on user point decomposition, and determine the number of output screens corresponding to each data output segment, including: Acquire the video collected at each point in the corresponding overlapping period based on each point collection unit corresponding to the immersive output point, and determine the target action attribute of the video collected at each point, wherein the target action attribute includes target action and no target action; Based on the preset output time, the video collected at each point corresponding to the target attribute of the target action is decomposed to obtain each data output segment; The data output quantity corresponding to each data output segment is determined as the output screen quantity.

5. The method according to claim 4, characterized in that Determine the target action attributes of the video captured at each point, wherein the target action attributes include target action and no target action, including: Determine each captured image frame constituting the captured video at each point, and perform image recognition on each captured image frame; Based on the recognition result, the collected image frames of each animal area corresponding to the same point collected video indicating the presence of any ornamental animal are aggregated into the same animal image group, and the video of the point collected video is updated based on the collected image frames located in the animal image group; In response to a video duration of a video captured at any point after the corresponding video is updated being greater than a preset time threshold, determining the target attribute of the video captured at the point as having a target action; Otherwise, the target attribute of the video captured at this point is determined to be no target action.

6. The method according to claim 5, characterized in that Based on the preset output time, the video collected at each point corresponding to the target attribute of the target action is decomposed to obtain each data output segment, including: Retrieving each captured image frame of the video captured at each point corresponding to the target attribute of having a target action, and determining each area center point of each animal area indicating each ornamental animal located in each captured image frame; Based on the video collected at each point, the center points of each area of ​​each collected image frame indicating the same ornamental animal with a neighboring relationship are divided into the same point group; Establishing each image coordinate system with the center point of each image of each captured image frame as the origin, and determining each region coordinate point corresponding to each region center point based on each image coordinate system; Determining the respective regional distances of the respective animal regions indicating the same ornamental animal based on the respective coordinate distances of the respective regional coordinate points corresponding to the same point group; In response to the regional distance being greater than the first preset distance, the captured image frames corresponding to the regional distance are aggregated into a moving image group corresponding to the point-captured video corresponding to each of the captured image frames; Determining the number of image frames corresponding to the moving image group based on each captured image frame in the same moving image group, and comparing the number of image frames with a preset number of frames determined based on a preset output time; In response to the number of image frames being greater than or equal to a preset number of frames, video fusion is performed on each of the captured image frames based on the preset number of frames, and point decomposition is performed on each of the obtained data output segments.

7. The method according to claim 6, characterized in that The method further comprises: In response to the number of image frames being less than a preset number of frames, determining the number of regions of each animal region in each captured image frame excluding the corresponding moving image group in the captured video at each point; In response to the presence of any region having a number greater than a preset number, adding captured image frames corresponding to the number of regions to the moving image group so that the captured image frames are connected to the last captured image frame in the moving image group, thereby obtaining a moving image group with an updated number of frames; In response to the image frame number of the moving image group corresponding to the frame number update being less than the preset frame number, each captured image frame whose corresponding moving distance is less than the first preset distance and greater than the second preset distance is added to the moving image group in sequence based on chronological order until the image frame number of the corresponding moving image group is greater than or equal to the preset frame number.

8. The method according to claim 6, characterized in that Perform point-wise decomposition on each data output segment obtained, including: determining respective image proportions corresponding to respective animal regions indicating respective ornamental animals located in output image frames constituting respective data output segments; In response to any image proportion being greater than a first preset proportion and less than a second preset proportion, determining an animal region corresponding to the image proportion as a first target region, and performing target extraction on the first target region; Amplifying the image of the extracted first target region based on a first magnification coefficient, and in response to the magnified first target region having an overlapping relationship with any animal region, reducing the image of the first target region based on a first reduction coefficient corresponding to the first magnification coefficient; In response to any image proportion being less than or equal to a second preset proportion, determining an animal region corresponding to the image proportion as a second target region, and performing target extraction on the second target region; Performing image magnification on the extracted second target area based on a second magnification coefficient, wherein the second magnification coefficient is greater than the first magnification coefficient; In response to the magnified second target area having an overlapping relationship with any animal area, the second target area is image-reduced based on a second reduction coefficient corresponding to the second magnification coefficient, and the second target area is determined as the first target area.

9. The method according to claim 6, characterized in that Interactively displaying each data output segment to a user terminal based on each data output screen includes: Sending each data output segment to each output display slot of each data output screen having an idle attribute; In response to the user completing wearing of any interactive unit, each data output segment is output and displayed, and based on the voice broadcast unit of the immersive data output area, the animal broadcast template of the immersive data collection area corresponding to each data output segment is voice broadcasted; In response to a user's interaction with any animal limb corresponding to an ornamental animal located in any data output segment based on the interaction unit, an interactive action corresponding to the animal limb with a preset number of interactive frames is generated based on the AI ​​module and interactively displayed to the user, and an interactive broadcast template corresponding to the interactive action is voice broadcasted based on the voice broadcast unit; Based on the preset interactive frame number, each captured image frame located at the end of the data output segment is removed.

10. An immersive AI scenic spot data output system, characterized in that: include: A determination module configured to determine, based on the immersive data collection area of ​​each first position, the first video segment corresponding to each user terminal that meets the immersive data output condition a decomposition module configured to, in response to any of the user terminals arriving at any immersive data output area corresponding to the second position, obtain the immersive output point of the first video segment and decompose it based on the user point to generate data output segments, and determine the number of output screens corresponding to the data output segments; The interactive module is configured to respond to the number of output screens corresponding to the data output screens with idle attributes located in the immersive data output area being greater than or equal to the number of output screens, and interactively display each data output segment to the user terminal based on each data output screen.

Citation Information

Patent Citations

  • Immersive video interaction method and device, equipment and storage medium

    CN113965802A

  • Immersive exhibition hall intelligent guide display method and system based on user behaviors

    CN120182488A

  • Novel real time physical reality immersive experiences having gamification of actions taken in physical reality

    US20190009177A1

  • Method and apparatus for managing immersive data

    US20190188828A1