Virtual reality video adjustment method and device, equipment and storage medium

By predicting users' emotional stimuli and generating interstitial videos, the problem of poor flexibility in VR video playback is solved, enabling personalized adjustments and improvements to the user experience.

CN120835187APending Publication Date: 2025-10-24XIAN UNIVIEW INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410486332.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-22
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

In existing VR experience centers, the scripts and plots of VR videos are fixed, which cannot take into account the experience of each user, resulting in poor playback flexibility and some users quitting midway due to excessive stimulation.

Method used

By predicting the emotional stimulus value of users during the viewing of VR videos, target moments are determined, and interstitial videos are generated to adjust the user's emotional stimulus value and inserted into specific moments in the VR video to adjust the user experience.

Benefits of technology

It improves the flexibility of VR video playback, enhances the experience for each user, and prevents users from quitting midway due to excessive or insufficient stimulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120835187A_ABST
    Figure CN120835187A_ABST
Patent Text Reader

Abstract

The invention provides a virtual reality video adjusting method and device, equipment and a storage medium, and the method comprises the steps: in a process that at least two users watch the same virtual reality video, adjusting the virtual reality video based on the user features of each user and a target video which is not played in the virtual reality video; predicting an emotional stimulation value corresponding to each moment when each user watches the target video; determining a target moment based on the emotional stimulus value corresponding to each moment when each user watches the target video; determining an inter-cut video based on the emotional stimulus value corresponding to at least one target user in the at least two users at the target moment, the inter-cut video being used for adjusting the emotional stimulus value of the at least one user; and inserting the inter-cut video after the target moment in the virtual reality video to obtain an adjusted virtual reality video. According to the invention, the flexibility of playing the virtual reality video can be improved, and the experience feeling of each user who participates in watching the same virtual reality video is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of video processing, and in particular to a virtual reality video adjustment method, device, equipment and storage medium. BACKGROUND

[0002] Virtual reality (VR) is a technology that uses computer technology to simulate and generate a simulated environment, and through special equipment enables users to immerse in it and interact with the virtual environment.

[0003] At present, in a VR experience hall, a user can experience an immersive experience by playing a pre-set VR video. When multiple users watch the same VR video in a group, the experience of each user will be different because each user can withstand different levels of stimulation, and even some users will quit halfway because the stimulation is too great. However, in the current VR experience hall, the plot of each VR video is fixed, and the experience of each user cannot be taken into account. Therefore, the flexibility of playing a VR video in the prior art is poor. SUMMARY

[0004] The present application provides a virtual reality video adjustment method, device, equipment and storage medium, which solves the defect that the flexibility of playing a VR video in the prior art is poor, and improves the flexibility of playing a VR video and the experience of each user watching the same VR video.

[0005] The present application provides a virtual reality video adjustment method, comprising:

[0006] During the process of at least two users watching the same virtual reality video, based on the user characteristics of each user and a target video that has not been played in the virtual reality video, the emotional stimulation value corresponding to each user at each time during the process of watching the target video is predicted;

[0007] Based on the emotional stimulation value corresponding to each user at each time during the process of watching the target video, a target time is determined;

[0008] Based on the emotional stimulation value corresponding to at least one target user in the at least two users at the target time, an intercalation video is determined, and the intercalation video is used to adjust the emotional stimulation value of the at least one target user;

[0009] The intercalation video is inserted after the target time in the virtual reality video, and an adjusted virtual reality video is obtained.

[0010] The virtual reality video adjustment method provided by the application comprises the following steps:

[0011] Extracting video description features of the target video;

[0012] Inputting the user features of each user and the video description features into an emotional stimulation prediction model to obtain emotional stimulation values corresponding to each user at each time point during the process of watching the target video output by the emotional stimulation prediction model, wherein the emotional stimulation prediction model is obtained by training an initial emotional stimulation prediction model based on user features of sample users and sample video description features of sample videos.

[0013] The virtual reality video adjustment method provided by the application comprises the following steps:

[0014] For each user, based on the emotional stimulation values corresponding to each user at each time point during the process of watching the target video, the time point at which the emotional stimulation value of the user reaches a preset maximum emotional stimulation value is determined.

[0015] The minimum time point among the time points at which all the users reach the preset maximum emotional stimulation value is determined as the target time point.

[0016] The virtual reality video adjustment method provided by the application comprises the following steps:

[0017] Based on the emotional stimulation values corresponding to each target user at the target time point, a target emotional stimulation value is determined.

[0018] Inputting the target emotional stimulation value into a video generation model to obtain the intercalation video output by the video generation model, wherein the video generation model is obtained by training an initial video generation model based on a first sample emotional stimulation value.

[0019] The virtual reality video adjustment method provided by the application comprises the following steps:

[0020] The average emotional stimulation value of the emotional stimulation values corresponding to all the target users at the target time point is determined as the target emotional stimulation value; or,

[0021] Determine a maximum emotional stimulus value of the emotional stimulus values corresponding to the target time of all the target users as the target emotional stimulus value.

[0022] According to the virtual reality video adjustment method provided by the application, the method further comprises:

[0023] For each preset video, determine a predicted emotional stimulus value of each target user when watching the preset video based on the video description feature of the preset video and the user feature of each target user.

[0024] Determine the difference between the predicted emotional stimulus value and the preset emotional stimulus value of each target user when watching each preset video.

[0025] Determine the preset video as the intercalation video, wherein the difference value corresponding to the same preset video of all target users in all preset videos is less than or equal to the preset value.

[0026] According to the virtual reality video adjustment method provided by the application, the method further comprises:

[0027] Determine the second sample emotional stimulus value of each sample time of the sample user in the process of watching the sample video.

[0028] If the change amount of the emotional stimulus value in the preset time length is greater than the preset change amount based on the second sample emotional stimulus value of each sample time, extract the sample video feature of the sample video segment in the preset time length.

[0029] Generate the preset video based on the sample video feature.

[0030] The application also provides a virtual reality video adjustment device, comprising:

[0031] A prediction module is configured to predict the emotional stimulus value corresponding to each time in the process of watching the target video of each user based on the user feature of each user and the target video not played in the same virtual reality video in the process of watching the same virtual reality video of at least two users.

[0032] A determination module is configured to determine the target time based on the emotional stimulus value corresponding to each time in the process of watching the target video of each user.

[0033] The determination module is further configured to determine an intercalation video based on the emotional stimulus value corresponding to the target time of at least one target user in the at least two users, and the intercalation video is used to adjust the emotional stimulus value of the at least one target user.

[0034] an adjusting module configured to adjust the inserted video in the virtual reality video based on the target time, to obtain an adjusted virtual reality video.

[0035] The application further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the virtual reality video adjusting method according to any one of the preceding embodiments when executing the program.

[0036] The application further provides a computer readable storage medium, which stores a computer program, wherein the computer program is executable on a processor to implement the virtual reality video adjusting method according to any one of the preceding embodiments.

[0037] The application further provides a computer program product, which comprises a computer program, wherein the computer program is executable on a processor to implement the virtual reality video adjusting method according to any one of the preceding embodiments.

[0038] The virtual reality video adjusting method, device, equipment and storage medium provided by the application can predict the emotional stimulation value corresponding to each time during the process of watching the target video for each user based on the user characteristics of each user and the target video not played in the VR video, determine the target time based on the emotional stimulation value corresponding to each time during the process of watching the target video for each user, and determine the inserted video based on the emotional stimulation value corresponding to the target time for at least one target user in the at least two users, wherein the determined inserted video is used to adjust the emotional stimulation value of each target user, and the determined inserted video is inserted after the target time in the VR video. Since the inserted video can be determined based on the emotional stimulation value corresponding to the target time for each target user when watching the target video, that is, the emotional stimulation value of each target user is considered when generating the inserted video, the emotional state of each target user can be adjusted through the inserted video, so that the acceptance degree of stimulation and the emotional state of each target user can be considered when playing the obtained adjusted VR video after inserting the inserted video in the VR video, thereby improving the flexibility of playing the VR video and the experience of each user participating in watching the same VR video. BRIEF DESCRIPTION OF DRAWINGS

[0039] In order to more clearly illustrate the technical solutions of the present application or the prior art, the following will briefly introduce the drawings needed in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.

[0040] Figure 1A flowchart of a virtual reality video adjustment method provided for an embodiment of the present application is shown in the figure;

[0041] Figure 2 An adjustment diagram of an emotional stimulation value provided for an embodiment of the present application is shown in the figure;

[0042] Figure 3 A diagram showing the time when three users respectively reach a preset maximum emotional stimulation value provided for an embodiment of the present application is shown in the figure;

[0043] Figure 4 Emotional stimulation values of three users respectively at t1 provided for an embodiment of the present application are shown in the figure;

[0044] Figure 5 An adjustment state diagram of an emotional stimulation value of a user provided for an embodiment of the present application is shown in the figure;

[0045] Figure 6 A structural diagram of a virtual reality video adjustment device provided for an embodiment of the present application is shown in the figure;

[0046] Figure 7 An example of a physical structure diagram of an electronic device is shown in the figure. DETAILED DESCRIPTION

[0047] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below with reference to the drawings in the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0048] In a VR experience hall, by playing a pre-set VR video, a user can achieve an immersive experience. However, in the current VR experience hall, the plot of each VR video is fixed. When multiple users watch the same VR video in a group, due to the difference in the acceptance ability of different users to stimulation, the experience of each user will be different. Some users can withstand a higher degree of stimulation, so they can try more stimulating scenes. On the contrary, some users have a weaker ability to withstand stimulation, and need to reduce the degree of stimulation, or even some individual users will quit halfway because the stimulation is too great. Therefore, the VR video in the prior art cannot take into account the emotional stimulation acceptance degree of each user, which makes the flexibility of playing the VR video in the prior art poor.

[0049] To solve the above problems, the embodiment of the present application provides a virtual reality video adjustment method. In the method, when at least two users watch the same VR, the emotional stimulation value of each user when watching the target video that has not been played can be predicted, and based on the emotional stimulation value of all or part of the target users at the target time, an interlude video is generated to adjust the emotional state of each target user through the interlude video. Since the interlude video can be determined based on the emotional stimulation value of each target user at the target time when watching the target video, that is, the emotional stimulation value of each target user is considered when generating the interlude video, the emotional state of each target user can be adjusted through the interlude video, so that when the adjusted VR video is played after the interlude video is inserted into the VR video, the acceptance degree of each target user to the stimulation and the emotional state can be considered, thereby not only improving the flexibility of playing the VR video, but also improving the experience of each user participating in watching the same VR video.

[0050] The following will be described in combination with Figures 1 to 5 The virtual reality video adjustment method provided by the embodiment of the present application is described. The embodiment of the present application can be applied to the scene where at least two users watch the same video and need to consider the emotional stimulation acceptance degree of each user to adjust the video, and can be particularly applied to the scene of adjusting the VR video. The execution subject of the method can be a terminal device, a computer, a server, a server cluster or a specially designed virtual reality video adjustment device, etc. electronic device, or a virtual reality video adjustment device provided in the electronic device, which can be realized by software, hardware or combination of both.

[0051] Figure 1 The flowchart of the virtual reality video adjustment method provided by the embodiment of the present application is shown in Figure 1 The method comprises:

[0052] Step 101: In the process of at least two users watching the same virtual reality video, based on the user characteristics of each user and the target video that has not been played in the virtual reality video, the emotional stimulation value of each user at each time during watching the target video is predicted.

[0053] In this step, the VR video is a video form that provides immersive visual and auditory experience for users through special devices (such as VR headsets, glasses, etc.). Users can directly touch the emotions by watching the VR video, so as to produce emotional stimulation value. The emotional stimulation value of the user can be used to measure the emotional stimulation degree of the user when watching the VR video.

[0054] In actual application, since the stimulation degree that each user can bear is different, even if at least two users form a team to watch the same VR video, the emotional stimulation value of each user can be different. In order to prevent the user with low stimulation bearing from quitting watching the VR video halfway, therefore, it is necessary to predict the emotional stimulation value of each user at each time during watching the target video based on the user characteristics of each user and the video description characteristics of the target video that has not been played, wherein the target video that has not been played in the VR video can be understood as the video that each user has not watched currently. Through the predicted emotional stimulation value, the emotional state of each user at each time during watching the target video can be obtained, and thus it can be judged which users can continue to watch the VR video and which users can not accept and quit halfway.

[0055] The user characteristics of the user can include identity information, face feature information or other information capable of uniquely identifying the user. The user characteristics can be manually input by the user or obtained through face recognition.

[0056] Step 102: determining a target time based on the emotional stimulation value corresponding to each time during watching the target video by each user.

[0057] In this step, after predicting the emotional stimulation value corresponding to each time during watching the target video by each user, the target time can be further determined based on the predicted emotional stimulation value. It should be understood that the target time is the time when the emotional stimulation value of all or part of the users participating in watching the VR video needs to be adjusted. The target time is a time after the current time, that is, a time that has not yet arrived.

[0058] Step 103: determining an inserted video based on the emotional stimulation value corresponding to the target time of at least one target user in the at least two users, wherein the inserted video is used to adjust the emotional stimulation value of the at least one target user.

[0059] In this step, the at least one target user can be all users in the at least two users, or part of the users in the at least two users that need to be adjusted in emotional stimulation value. The users that need to be adjusted in emotional stimulation value can include at least one of the following users: a user whose predicted emotional stimulation value during watching the target video exceeds a preset emotional stimulation value, a user whose emotional stimulation value is less than the preset emotional stimulation value, and a user whose emotional stimulation value fluctuation value is large. The preset emotional stimulation value can be understood as the best ideal emotional state of each user during watching the VR video, that is, the value to which the emotional stimulation values of the three users are adjusted.

[0060] Taking three users participating in watching the same VR video as an example, it is assumed that the predicted emotional stimulation values of the three users when watching the target video are S1, S2 and S3 respectively, and the preset emotional stimulation value is S0. The difference between the emotional stimulation value of each user and the preset emotional stimulation value S0, that is, ΔSi = S0-Si, is calculated, where Δsi represents the difference between the preset emotional stimulation value S0 and the emotional stimulation value Si of the i-th user. The user corresponding to the maximum Δsi can be determined as the target user, or the user corresponding to the non-zero Δsi can be determined as the target user, and the like.

[0061] According to the predicted emotional stimulation value of each user at each time when watching the target video, the emotional stimulation value of each target user at the target time can be found. By comprehensively considering the emotional stimulation value of each target user at the target time, an interlude video can be generated, which may be, for example, a video containing light music, or a video containing beautiful scenery, and the like. Since the emotional stimulation value of each target user is considered when generating the interlude video, the purpose of adjusting the emotional stimulation value of each target user can be achieved through the playing of the interlude video. For example, assuming that the emotional stimulation value of the target user is high, the interlude video can be played to reduce the emotional stimulation value of the target user. For another example, assuming that the emotional stimulation value of the target user is low, the interlude video can be played to improve the emotional stimulation value of the target user. In this way, the situation that some users quit halfway during watching the VR video due to unbearable stimulation or boredom of the VR video can be prevented.

[0062] Step 104: inserting the interlude video after the target time in the virtual reality video to obtain an adjusted virtual reality video.

[0063] In this step, after the interlude video is generated, the interlude video can be inserted after the target time in the VR video to adjust the VR video and obtain an adjusted VR video. Since the target time is the time when the emotional stimulation value of the target user needs to be adjusted, the interlude video is inserted after the target time, and when the interlude video is played, the target user will adjust the emotional stimulation value during watching the interlude video, thereby achieving the purpose of adjusting the emotional state of the target user. It should be understood that adding the interlude video after the target time can adjust the plot development in the virtual environment in the VR video, such as adjusting the elements of events, character behaviors and sound effects in the virtual environment, so as to adjust the current emotional state of the user watching the VR video and improve the emotional experience of each user.

[0064] In specific applications, during the VR video playback process, the emotional stimulation value of each user when watching the unplayed video can be predicted in real time, and an interstitial video can be generated in accordance with the method described in the embodiment of the present invention to adjust the VR video. In this way, the purpose of dynamically adjusting the VR video and dynamically adjusting the emotional stimulation value of each target user can be achieved, thereby adjusting the emotional state of each target user. When at least two users watch the same VR video synchronously, each user can be in the best emotional state, avoiding the phenomenon that some users quit the experience midway due to not being able to accept the stimulation level of the VR video, thereby improving the overall user experience of the team.

[0065] For example, Figure 2 The schematic diagram of adjusting the emotional stimulation value provided by the embodiment of the present invention is as follows: Figure 2 As shown, when the emotional stimulation value of a user is greater than the preset maximum emotional stimulation value SH, that is, when it is within the p1 segment, it indicates that the user is currently in a state of excessive excitement or fear. If the emotional stimulation value is less than the preset minimum emotional stimulation value SL, that is, when it is within the p2 segment, it indicates that the user is currently in a state of excessive depression. Therefore, when it is predicted that the emotional stimulation value of a user is in the p1 segment or the p2 segment, it is necessary to dynamically insert an interstitial video to adjust the user's emotional stimulation value so that the adjusted emotional stimulation value is within the p3 segment, that is, to ensure that the emotional stimulation values ​​of all users participating in watching the same VR video are controlled between the preset maximum emotional stimulation value SH and the preset minimum emotional stimulation value SL. Among them, the interstitial video can include, for example, a video for emotional soothing or a game video.

[0066] The virtual reality video adjusting method provided by the embodiments of the present application can predict the emotional stimulation value corresponding to each time point in the process of each user watching a target video based on the user features of each user and the target video that has not been played in the process of at least two users watching the same VR video, determine a target time point based on the emotional stimulation value corresponding to each time point in the process of each user watching the target video, and determine an inserted video based on the emotional stimulation value corresponding to the target time point of at least one target user in the at least two users, wherein the determined inserted video is used to adjust the emotional stimulation value of each target user, and the determined inserted video is inserted after the target time point in the VR video. Since the inserted video can be determined based on the emotional stimulation value corresponding to the target time point of each target user when watching the target video, that is, the emotional stimulation value of each target user is considered when generating the inserted video, the emotional state of each target user can be adjusted through the inserted video, so that the acceptance degree of stimulation and the emotional state of each target user can be taken into account when playing the adjusted VR video obtained after inserting the inserted video in the VR video, thereby not only improving the flexibility of playing the VR video, but also improving the experience of each user participating in watching the same VR video.

[0067] For example, based on the above embodiments, when predicting the emotional stimulation value corresponding to each time point in the process of each user watching a target video based on the user features of each user and the target video that has not been played in the virtual reality video, the video description features of the target video can be extracted, and the user features of each user and the video description features are input into an emotional stimulation prediction model to obtain the emotional stimulation value corresponding to each time point in the process of each user watching the target video output by the emotional stimulation prediction model. The emotional stimulation prediction model is obtained by training an initial emotional stimulation prediction model based on the user features of sample users and sample video description features of sample videos.

[0068] Specifically, for a target video that has not been played yet, the video description features of the target video can be extracted, which can be used to represent the plot description information of the target video. Through the video description features, the plot content of the target video can be determined.

[0069] For each user, the user features of the user and the extracted video description features can be input into a pre-trained emotional stimulation prediction model, and the emotional stimulation value of the user predicted by the emotional stimulation prediction model at each time point in the process of watching the target video can be obtained.

[0070] When training the emotional stimulation prediction model, the user features of sample users and the sample video description features of sample videos are first needed to be collected, wherein the sample videos include historical VR videos watched by the sample users before, and can also include the VR videos watched this time.

[0071] In addition, the real emotional stimulus value of the sample user at each time when watching the sample video needs to be determined as a label during model training. Specifically, the real emotional stimulus value of each sample user when watching the sample video can be detected in real time by establishing an emotional recognition sensor system. For example, the facial expression of the sample user can be obtained by collecting facial images of the user and recognizing the facial images through facial expression recognition technology. In addition, eye tracking data of the sample user and behavior data of the sample user can also be recognized, and physiological signals of the sample user can be detected through a physiological data monitoring sensor. Through the obtained facial expression, eye tracking data, behavior data and physiological signals, etc., the real emotional stimulus value of the sample user when watching the sample video can be determined. The physiological signals may, for example, include blood pressure, heart rate, respiratory rate or galvanic skin response, etc.

[0072] Further, the initial emotional stimulus prediction model may, for example, be a deep neural network model, which can include an input layer, a hidden layer and an output layer. The input layer is used to input the user features of the sample user and the sample video description features of the sample video. The hidden layer is a multi-layer neural network structure, which is used to learn the complex relationship between the features. The output layer is used to predict the predicted emotional stimulus value of the sample user when watching the sample video.

[0073] In addition, a loss function needs to be defined, for example, a mean square error loss function is defined as follows:

[0074]

[0075] wherein N represents the number of sample videos, S i represents the real emotional stimulus value, S i ′ represents the predicted emotional stimulus value.

[0076] After the user features of the sample user and the sample video description features of the sample video are input into the initial emotional stimulus prediction model to obtain the predicted emotional stimulus value, the loss information can be determined according to the above loss function, and the neural network parameters of the initial emotional stimulus prediction model are updated through the back propagation algorithm. By repeatedly executing the above process, the final obtained model is determined as the emotional stimulus prediction model until the loss function is minimized or the iteration number reaches the preset number.

[0077] In this embodiment, after the video description features of the target video are extracted, the user features of each user and the extracted video description features are input into the emotional stimulus prediction model, and the emotional stimulus prediction model can quickly and accurately predict the emotional stimulus value of each user at each time when watching the target video.

[0078] For example, based on the above embodiments, when the target time is determined based on the emotional stimulation values corresponding to each time during the process of watching the target video by each user, for each user, the time when the emotional stimulation value of the user reaches the preset maximum emotional stimulation value can be determined based on the emotional stimulation values corresponding to each time during the process of watching the target video by the user, and the minimum time among the times when all users reach the preset maximum emotional stimulation value is determined as the target time.

[0079] Specifically, for each user, the emotional stimulation value corresponding to each time during the process of watching the target video by the user can be determined, so that the time when the emotional stimulation value of the user reaches the preset maximum emotional stimulation value can be further determined. The preset maximum emotional stimulation value can represent the maximum emotional stimulation degree that the user can withstand. Since each user can withstand different stimulation degrees, different preset maximum emotional stimulation values can be set for each user. Of course, the same preset maximum emotional stimulation value can also be set, for example, the preset maximum emotional stimulation value can be set according to the user who can withstand the lowest stimulation degree.

[0080] Taking the example that each user corresponds to a different preset maximum emotional stimulation value, the time when the emotional stimulation value of each user reaches its own corresponding preset maximum emotional stimulation value can be determined based on the emotional stimulation values corresponding to each time during the process of watching the target video by each user.

[0081] Further, in order to avoid the phenomenon that some users quit halfway because they cannot withstand the stimulation and no longer continue to participate in watching the target video, it is necessary to insert an interlude video at the first time when the preset maximum emotional stimulation value is reached, so as to timely adjust the emotional stimulation value of the first user who reaches the preset maximum emotional stimulation value. In summary, the minimum time among the times when all users reach the preset maximum emotional stimulation value can be determined as the target time.

[0082] It should be noted that there can be some users whose emotional stimulation values do not reach the preset maximum emotional stimulation value at all times. At this time, it indicates that these users can continue to watch the target video without adjusting the emotional state, so when the target time is determined, the emotional stimulation values of these users can not be considered.

[0083] Taking the example that three users are watching the same VR video at the same time, Figure 3 The schematic diagram of the times when the three users respectively reach the preset maximum emotional stimulation value provided by the embodiments of the present application is shown in FIG. 1. Figure 3 As shown in FIG. 1, user 1 reaches its own corresponding preset maximum emotional stimulation value at t1, user 2 reaches its own corresponding preset maximum emotional stimulation value at t2, and user 3 reaches its own corresponding preset maximum emotional stimulation value at t3, wherein t1 is less than t2, and t2 is less than t3.

[0084] After predicting the time when each user reaches the preset maximum emotional stimulation value when watching the target video, in order to prevent the case that a user cannot accept the subsequent plot and exits halfway, therefore, the minimum time when all users reach the preset maximum emotional stimulation value can be determined as the target time, such as determining t1 as the target time, so that when the first user 1 who reaches the preset maximum emotional stimulation value cannot accept the subsequent plot, the VR video is adjusted in time to achieve the purpose of adjusting the emotional stimulation value of the user 1.

[0085] Figure 4 The emotional stimulation values of the three users at t1 are provided in the embodiment of the application, as shown in the table. Figure 4 The emotional stimulation value of the user 1 at t1 is s1, the emotional stimulation value of the user 2 at t1 is s2, and the emotional stimulation value of the user 3 at t1 is s3. It can be known from the table that Figure 4 At t1, the emotional stimulation value of the user 1 has reached the preset maximum emotional stimulation value.

[0086] In the embodiment, the minimum time when all users reach the preset maximum emotional stimulation value can be determined as the target time, so that the emotional state of the user who first reaches the preset maximum emotional stimulation value can be adjusted in time, and the case that the user exits halfway when watching the VR video due to unbearable stimulation is avoided, not only the adjustment of the VR video is more timely, but also the experience of each user participating in watching the same VR video can be improved.

[0087] For example, in a possible implementation, when the intercalation video is determined based on the emotional stimulation value corresponding to the target time of at least one target user in at least two users, the emotional stimulation value corresponding to the target time of each target user can be determined, and then the target emotional stimulation value is input into a video generation model to obtain the intercalation video output by the video generation model, and the video generation model is obtained by training the initial video generation model based on the first sample emotional stimulation value.

[0088] Specifically, the target emotional stimulus value can be determined based on the emotional stimulus value corresponding to each target user at the target time. For example, the average emotional stimulus value of the emotional stimulus values corresponding to all target users at the target time can be determined as the target emotional stimulus value. In this way, the target emotional stimulus value takes into account the emotional stimulus values of all target users at the target time, and thus the generated interlude video can serve to adjust the emotional state of all target users. Alternatively, the maximum emotional stimulus value among the emotional stimulus values corresponding to all target users at the target time can be determined as the target emotional stimulus value. In this way, the interlude video generated based on the target emotional stimulus value can serve to adjust the emotional state of the target user corresponding to the maximum emotional stimulus value, and thus the interlude video is more targeted.

[0089] After the target emotional stimulus value is determined, the target emotional stimulus value can be input into the pre-trained video generation model to obtain an interlude video output by the video generation model. When training the video generation model, the first sample emotional stimulus values of different sample users can be collected, and a real video corresponding to the first sample emotional stimulus values is determined as a label for model training. When the sample users watch the real video, the corresponding first sample emotional stimulus values tend to be the pre-set ideal preset emotional stimulus value, i.e., the sample users adjust the emotional stimulus value to an ideal state when watching the real video.

[0090] After the first sample emotional stimulus value is input into the initial video generation model, a predicted video is obtained. The loss information is determined based on the predicted video and the real video, and the model parameters of the initial video generation model are adjusted based on the loss information. The above process is repeatedly performed until the model converges or the number of iterations reaches a pre-set number. The final obtained model is determined as the trained video generation model.

[0091] In this embodiment, the target emotional stimulus value can be input into the video generation model, and the interlude video can be quickly and accurately generated by the video generation model.

[0092] In another possible implementation, when generating the interlude video, the predicted emotional stimulus value of each target user when watching each preset video can be determined based on the video description features of the preset video and the user features of each target user. After the difference between the predicted emotional stimulus value of each target user when watching each preset video and the preset emotional stimulus value is determined, the preset video in which the difference between the predicted emotional stimulus value of each target user corresponding to the same preset video and the preset value is less than or equal to the preset value is determined as the interlude video.

[0093] Specifically, a plurality of preset videos are pre-stored in the video library, for each preset video, a video description feature of the preset video can be extracted, and the extracted video description feature and the user feature of each target user are input into the emotion prediction model, so that the predicted emotional stimulation value of each target user when watching the preset video output by the emotion prediction model can be obtained. It should be understood that the predicted emotional stimulation value of the target user at each moment when watching the preset video can be output by the emotion prediction model, and the average value of these predicted emotional stimulation values can be used as the final predicted emotional stimulation value.

[0094] In addition, the preset emotional stimulation value can be understood as an ideal emotional stimulation value, that is, a value that is pre-set and intended to be finally adjusted to the emotional stimulation value of each user. Therefore, in order to find the most suitable preset video from a plurality of preset videos, it is necessary to determine the difference between the predicted emotional stimulation value of each target user when watching each preset video and the preset emotional stimulation value.

[0095] Further, from all the preset videos, the preset video corresponding to the difference value of all users for the same preset video is found to be less than or equal to the preset value, and the preset video is determined as the interlude video. Since the difference between the predicted emotional stimulation value of all target users when watching the interlude video and the preset emotional stimulation value is less than or equal to the preset value, it indicates that the emotional state of each target user when watching the interlude video is in a relatively ideal state.

[0096] Among them, the preset video includes a video that can relieve emotions, such as a video containing beautiful scenery and a video containing light music, etc., to help users calm down. The preset video can also include a video of role emotional interaction, such as expressing understanding, support or comfort according to the emotional state of the user, so as to guide the user's emotion to develop in a positive direction. In addition, the preset video can also be a preset video customized according to the historical emotional stimulation value and feedback information of the user, which is targeted or personalized to better meet the emotional needs of individual differences of the user.

[0097] Figure 5 The adjustment state diagram of the emotional stimulation value of the user provided by the embodiment of the present application is shown in Figure 5 As shown, by selecting a suitable preset video as an interlude video, after playing the interlude video, the emotional stimulation value of each user watching the same VR video can be adjusted to a suitable state, such as adjusting to the preset emotional stimulation value.

[0098] In the embodiment, the predicted emotional stimulation value of each target user when watching each preset video can be predicted, and from all the preset videos, the preset video whose predicted emotional stimulation value of all users when watching the same preset video is close to the preset emotional stimulation value is selected as the inserted video. The inserted video determined by the above-mentioned manner can ensure that the emotional stimulation value of each target user tends to the preset emotional stimulation value, so that the accuracy of the determined inserted video is higher, and the effect of adjusting the emotional stimulation value of each target user by the inserted video is better.

[0099] For example, on the basis of the above-mentioned embodiment, when determining each preset video in the video library, the second sample emotional stimulation value of the sample user at each sample moment in the process of watching the sample video can be determined. When the change amount of the emotional stimulation value in the preset time period is greater than the preset change amount, the sample video features of the sample video segment in the preset time period are extracted, so as to generate the preset video based on the sample video features.

[0100] Specifically, the user features of the sample user and the video description features of the sample video can be input into the emotional stimulation prediction model, and the second sample emotional stimulation value of the sample user at each sample moment in the process of watching the sample video output by the emotional stimulation prediction model can be obtained. Based on the second sample emotional stimulation value at each sample moment, it can be judged whether the sample video contains a video segment that makes the emotional fluctuation of the sample user larger, that is, whether the change amount of the emotional stimulation value of the sample user in a certain preset time period is greater than the preset change amount. If so, the sample video features of the sample video segment in the preset time period are extracted. The sample video features can be understood as the features that can cause the emotion of the sample user to fluctuate greatly in the sample video segment, for example, the features that can cause the emotional stimulation value of the sample user to decrease or the features that can cause the emotional stimulation value of the sample user to increase. Further, the preset video can be generated based on the extracted sample video features. By the above-mentioned manner, a plurality of preset videos can be generated, so as to store these preset videos in the video library. When determining the inserted video subsequently, a suitable preset video can be selected from the video library for playing to adjust the emotional state of the user.

[0101] In the embodiment, the second sample emotional stimulus value of the sample user at each sample moment during watching the sample video can be determined, and based on the second sample emotional stimulus value, a sample video segment capable of causing a relatively large fluctuation of the emotion of the sample user can be determined. After the sample video feature of the sample video segment is extracted, the preset video can be generated based on the sample video feature. The preset video generated by the above manner is more abundant, and these preset videos are the videos that can really cause the emotion fluctuation of the sample user, so the accuracy of the inserted video determined based on these preset videos is higher.

[0102] The virtual reality video adjustment device provided by the present application is described below. The virtual reality video adjustment device described below can be referred to in correspondence with the virtual reality video adjustment method described above.

[0103] Figure 6 The structural schematic diagram of the virtual reality video adjustment device provided by the embodiment of the present application is shown in Figure 6 The virtual reality video adjustment device 600 includes:

[0104] The prediction module 601 is configured to, during watching of the same virtual reality video by at least two users, predict an emotional stimulus value corresponding to each moment during watching of a target video by each of the users based on a user feature of each of the users and the target video not played in the virtual reality video.

[0105] The determination module 602 is configured to determine a target moment based on the emotional stimulus value corresponding to each moment during watching of the target video by each of the users.

[0106] The determination module 602 is further configured to determine an inserted video based on the emotional stimulus value corresponding to the target moment of at least one target user in the at least two users, the inserted video being used to adjust the emotional stimulus value of the at least one user.

[0107] The adjustment module 603 is configured to adjust the inserted video in the virtual reality video based on the target moment, to obtain an adjusted virtual reality video.

[0108] In an example embodiment, the prediction module 601 is specifically configured to:

[0109] extract a video description feature of the target video;

[0110] input the user features of each of the users and the video description features into an emotional stimulus prediction model to obtain emotional stimulus values of each of the users at each time during watching of the target video output by the emotional stimulus prediction model, the emotional stimulus prediction model being obtained by training an initial emotional stimulus prediction model based on user features of sample users and sample video description features of sample videos.

[0111] In an example embodiment, the determining module 602 is specifically configured to:

[0112] For each of the users, based on the emotional stimulus value of the user at each time during watching of the target video, determine a time when the emotional stimulus value of the user reaches a preset maximum emotional stimulus value;

[0113] Determine the minimum time among the times when all the users reach the preset maximum emotional stimulus value as the target time.

[0114] In an example embodiment, the determining module 602 is specifically configured to:

[0115] Determine a target emotional stimulus value based on the emotional stimulus value of each of the target users at the target time;

[0116] Input the target emotional stimulus value into a video generation model to obtain the inserted video output by the video generation model, the video generation model being obtained by training an initial video generation model based on a first sample emotional stimulus value.

[0117] In an example embodiment, the determining module 602 is specifically configured to:

[0118] Determine an average emotional stimulus value of the emotional stimulus values of all the target users at the target time as the target emotional stimulus value; or

[0119] Determine a maximum emotional stimulus value in the emotional stimulus values of all the target users at the target time as the target emotional stimulus value.

[0120] In an example embodiment, the determining module 602 is specifically configured to:

[0121] For each of the preset videos, determine a predicted emotional stimulus value of each of the target users when watching the preset video based on a video description feature of the preset video and user features of each of the target users;

[0122] Determine a difference between the predicted emotional stimulus value of each of the target users when watching each of the preset videos and a preset emotional stimulus value;

[0123] Among all the preset videos, a preset video whose corresponding difference values ​​for the same preset video by all target users are less than or equal to the preset value is determined as the interstitial video.

[0124] In an exemplary embodiment, the apparatus further comprises an extraction module and a generation module, wherein:

[0125] The determination module 602 is further configured to determine a second sample emotional stimulation value at each sample moment in the process of the sample user watching the sample video;

[0126] An extraction module, configured to extract sample video features of the sample video clip within the preset time length when it is determined based on the second sample emotional stimulation value at each of the sample moments that a change in the emotional stimulation value within the preset time length is greater than a preset change amount;

[0127] A generating module is used to generate the preset video based on the sample video features.

[0128] The device of this embodiment can be used to execute the method of any embodiment in the embodiment of the virtual reality video adjustment method side. Its specific implementation process and technical effects are similar to those in the embodiment of the virtual reality video adjustment method side. For details, please refer to the detailed description in the embodiment of the virtual reality video adjustment method side, and will not be repeated here.

[0129] Figure 7 An example of a physical structure diagram of an electronic device is shown below. Figure 7 As shown, the electronic device may include: a processor 710, a communications interface 720, a memory 730, and a communications bus 740, wherein the processor 710, the communications interface 720, and the memory 730 communicate with each other via the communications bus 740. The processor 710 may invoke logic instructions in the memory 730 to execute a virtual reality video adjustment method, the method comprising: while at least two users are viewing the same virtual reality video, predicting an emotional stimulation value corresponding to each user at each moment while viewing the target video based on user characteristics of each user and a target video not played in the virtual reality video; determining a target moment based on the emotional stimulation value corresponding to each user at each moment while viewing the target video; determining an interstitial video based on the emotional stimulation value corresponding to at least one target user among the at least two users at the target moment, the interstitial video being used to adjust the emotional stimulation value of the at least one target user; and inserting the interstitial video into the virtual reality video after the target moment to obtain an adjusted virtual reality video.

[0130] In addition, the logic instructions in the memory 730 described above can be implemented in the form of software functional units and sold or used as independent products, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for making a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0131] In another aspect, the present application also provides a computer program product, which comprises a computer program, the computer program can be stored on a computer readable storage medium, and the computer program can be executed by a processor to enable a computer to execute the virtual reality video adjustment method provided by the above-mentioned method. The method comprises: during the process that at least two users watch the same virtual reality video, predicting the emotional stimulation value corresponding to each user at each time during the process that the user watches the target video based on the user characteristics of each user and the target video not played in the virtual reality video; determining a target time based on the emotional stimulation value corresponding to each user at each time during the process that the user watches the target video; determining an intercalated video based on the emotional stimulation value corresponding to at least one target user in the at least two users at the target time, the intercalated video being used to adjust the emotional stimulation value of the at least one target user; and inserting the intercalated video after the target time in the virtual reality video to obtain an adjusted virtual reality video.

[0132] In yet another aspect, the present application also provides a computer readable storage medium having stored thereon a computer program, which, when executed by a processor, implements the virtual reality video adjustment method provided by the above method, and the method comprises: in a process in which at least two users watch a same virtual reality video, predicting, based on user features of each of the users and a target video that is not played in the virtual reality video, an emotional stimulation value corresponding to each of the users at each time point in a process of watching the target video; determining a target time point based on the emotional stimulation value corresponding to each of the users at each time point in the process of watching the target video; determining an interlude video based on an emotional stimulation value corresponding to at least one target user among the at least two users at the target time point, the interlude video being used to adjust the emotional stimulation value of the at least one target user; and inserting the interlude video after the target time point in the virtual reality video to obtain an adjusted virtual reality video.

[0133] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., can be located in one place or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment scheme according to actual needs. Those skilled in the art can understand and implement without creative labor.

[0134] From the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be realized by means of software and necessary general hardware platforms, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the method described in each embodiment or some parts of the embodiment.

[0135] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A virtual reality video adjustment method, characterized by, The method comprises the following steps: During the process that at least two users watch the same virtual reality video, based on the user characteristics of each of the users and a target video that has not been played in the virtual reality video, a corresponding emotional stimulus value of each of the users at each time during the process of watching the target video is predicted; Based on the corresponding emotional stimulus value of each of the users at each time during the process of watching the target video, a target time is determined; Based on the corresponding emotional stimulus value of at least one target user in the at least two users at the target time, an interlude video is determined, the interlude video being used to adjust the emotional stimulus value of the at least one target user; The interlude video is inserted after the target time in the virtual reality video, so that an adjusted virtual reality video is obtained.

2. The virtual reality video adjustment method of claim 1, wherein, The step of predicting the corresponding emotional stimulus value of each of the users at each time during the process of watching the target video based on the user characteristics of each of the users and the target video that has not been played in the virtual reality video comprises the following steps: Video description characteristics of the target video are extracted; The user characteristics of each of the users and the video description characteristics are input into an emotional stimulus prediction model, so that the corresponding emotional stimulus value of each of the users at each time during the process of watching the target video output by the emotional stimulus prediction model is obtained, the emotional stimulus prediction model being obtained by training an initial emotional stimulus prediction model based on sample user characteristics of sample users and sample video description characteristics of sample videos.

3. The virtual reality video adjustment method of claim 1, wherein, The step of determining the target time based on the corresponding emotional stimulus value of each of the users at each time during the process of watching the target video comprises the following steps: For each of the users, based on the corresponding emotional stimulus value of the user at each time during the process of watching the target video, a time when the emotional stimulus value of the user reaches a preset maximum emotional stimulus value is determined; The minimum time among the times when the emotional stimulus values of all the users reach the preset maximum emotional stimulus value is determined as the target time.

4. The virtual reality video adjustment method of any of claims 1-3, wherein, The step of determining the interlude video based on the corresponding emotional stimulus value of at least one target user in the at least two users at the target time comprises the following steps: Based on the corresponding emotional stimulus value of each of the target users at the target time, a target emotional stimulus value is determined; The target emotional stimulus value is input into a video generation model, so that the interlude video output by the video generation model is obtained, the video generation model being obtained by training an initial video generation model based on a first sample emotional stimulus value.

5. The virtual reality video adjustment method of claim 4, wherein, The step of determining the target emotional stimulus value based on the corresponding emotional stimulus value of each of the target users at the target time comprises the following steps: An average emotional stimulus value of the corresponding emotional stimulus values of all the target users at the target time is determined as the target emotional stimulus value; or A maximum emotional stimulus value among the corresponding emotional stimulus values of all the target users at the target time is determined as the target emotional stimulus value.

6. The virtual reality video adjustment method of any of claims 1-3, wherein, The step of determining the interlude video based on the corresponding emotional stimulus value of at least one target user in the at least two users at the target time comprises the following steps: For each preset video, based on video description characteristics of the preset video and the user characteristics of each of the target users, a predicted emotional stimulus value of each of the target users when watching the preset video is determined; determining a difference between a predicted emotional stimulus value of each of the target users and a preset emotional stimulus value when each of the target users watches each of the preset videos; determining, as the interlude video, a preset video in which a difference between a preset video corresponding to a same preset video and a preset value is less than or equal to a preset value for all target users in all preset videos.

7. The virtual reality video adjustment method of claim 6, wherein, The method further comprises: determining a second sample emotional stimulus value of a sample user at each sample time during watching of a sample video; extracting a sample video feature of a sample video segment in a preset time period in which a change in the emotional stimulus value is greater than a preset change amount based on the second sample emotional stimulus value at each sample time; generating the preset video based on the sample video feature.

8. A virtual reality video adjustment method, characterized by, comprises: a prediction module configured to predict an emotional stimulus value corresponding to each time during watching of a target video by each of at least two users based on a user feature of each of the users and the target video that is not played in a virtual reality video; a determination module configured to determine a target time based on the emotional stimulus value corresponding to each time during watching of the target video by each of the users; the determination module is further configured to determine an interlude video based on the emotional stimulus value corresponding to the target time of at least one target user in the at least two users, the interlude video being used to adjust the emotional stimulus value of the at least one user; an adjustment module configured to adjust the interlude video in the virtual reality video based on the target time to obtain an adjusted virtual reality video.

9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the virtual reality video adjustment method according to any one of claims 1 to 7 when executing the program.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program implements the virtual reality video adjustment method according to any one of claims 1 to 7 when executed by the processor.

Citation Information

Patent Citations

  • Playing method and device of virtual reality scene contents

    CN106249903A

  • Video emotion classification method and system fusing electroencephalogram and stimulation source information

    CN113095428A