A panoramic video transmission system and method based on line-of-sight guidance
By introducing a line of sight guidance module in the panoramic video transmission system, the user's viewing line of sight is optimized, and combined with multicast technology, the problem of huge overhead of panoramic video transmission traffic is solved, and effective traffic reduction and user preferences are achieved.
Patent Information
- Application Number
- CN202211276800.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-18
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-10-18
Smart Images

Figure CN115665446B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of panoramic video transmission, and particularly relates to a panoramic video transmission system and method based on line-of-sight guidance. Background Art
[0002] Virtual Reality (VR) panoramic video transmission, as an important application scenario of 5G and next-generation mobile communications, has the value of promoting smart life and has the characteristics of "high bandwidth, large storage, low latency, and complex visual adaptability". Panoramic video transmission will bring a relatively high traffic burden to the communication system. The chunk transmission method can effectively reduce the system transmission traffic. It divides the panoramic video frame into a series of video chunks and only transmits the frames within the user's Field of View (FOV). The proactive chunk transmission method predicts the user's future FOV and pre-transmits the chunks to be viewed, which helps to reduce the latency of panoramic video transmission [Reference 1: X. Wei, C. Yang, and S. Han, "Prediction, communication, and computing duration optimization for VR video streaming," IEEE Trans. Commun., vol. 69, no. 3, pp. 1947–1959, Nov. 2020.]. Multicast technology is an effective method to reduce the transmission traffic of video-on-demand services [Reference 2: S. Han, X. Tan, K. Qi, C. Yang, A. F. Molisch, Y. Lu, J. Zheng, and Y. Li, "Rethinking the gain of multicasting and proactive caching for VoD service," IEEE Wirel. Commun., vol. 27, no. 5, pp. 133–139, Oct. 2020.]. For the chunk transmission of panoramic videos, multicast can simultaneously serve multiple requests for the same chunk in the same video frame.
[0003] The traditional multicast transmission system based on panoramic video chunks consists of a server, a communication link, and a client. At the server side, the panoramic video frame is divided into several chunks after projection. The client collects and reports the user's viewing viewpoints to the server. The server calculates the user's field of view based on the received user viewpoints and extracts the video chunks within the FOV. If multiple users view the same chunk simultaneously, multicast transmission is used; otherwise, the chunk is sent to each user individually.
[0004] Although the multicast technology based on video chunking has the advantage of saving traffic, its performance depends on the concentration of users' FOV. If the concentration of FOV that users watch is very low, the advantages of this scheme will be significantly weakened. Summary of the Invention
[0005] In the face of the huge overhead problem of panoramic video transmission traffic, the present invention provides a panoramic video transmission system and method based on line-of-sight guidance. By introducing line-of-sight guidance in the panoramic video transmission system, the concentration of users' viewing fields of view is improved by guiding users' viewing lines of sight, the multicast transmission opportunity is increased, the panoramic video transmission traffic is reduced, and at the same time, users' viewing preferences are maintained.
[0006] A panoramic video transmission system based on line-of-sight guidance provided by the present invention improves the multicast transmission system based on panoramic video chunking. A line-of-sight guidance point optimization and distribution module is newly added in the server, and a line-of-sight guidance result presentation module is newly added in the client. The client obtains the user's current viewpoint in real time and sends it to the server.
[0007] The line-of-sight guidance point optimization and distribution module of the server establishes a statistical model of the user's viewpoint position after line-of-sight guidance according to the user's current viewpoint, the probability that the user accepts line-of-sight guidance, and the line-of-sight guidance habit of the user after receiving the guidance, and calculates the average transmission traffic of the system after line-of-sight guidance, as well as the degree of deviation of the viewpoint position after guidance from the user's viewpoint position when not guided. Then, it optimizes the line-of-sight guidance position by minimizing the weighted sum of the average transmission traffic of the system and the user's viewpoint offset, and sends the optimized line-of-sight guidance position to the client. The line-of-sight guidance result presentation module of the client displays the line-of-sight guidance position obtained from the server to the user for the user to select.
[0008] Correspondingly, a panoramic video transmission method based on line-of-sight guidance provided by the present invention includes the following steps:
[0009] Step 1, the client obtains the user's current viewpoint in real time and sends it to the server;
[0010] Step 2, the server calculates the viewpoint position of the user after the current line-of-sight guidance, including:
[0011] First, determine the line-of-sight guidance position z. Let the current viewpoint of user k be p k , and user k moves towards z with probability φ k to accept the guidance. Let the viewpoint after movement be If the guidance is not accepted, the user's viewpoint will change to the predicted viewpoint according to p k Both the viewing point and the line-of-sight guiding position are plural, and the real part and the imaginary part respectively correspond to the longitude coordinate and the latitude coordinate of the position; then, according to the probability that the user accepts the line-of-sight guidance and the line-of-sight guidance habit after the user receives the guidance, a statistical model of the user's viewing point position after the line-of-sight guidance is established;
[0012] Among them, the spherical screen of the panoramic video is projected onto a plane and cut into M blocks and N blocks in terms of longitude and latitude. A vector w(p k ) is used to record the blocks transmitted to the user. The transmitted blocks need to cover the user's field of view. When the value of each dimension of the vector w(p k ) takes 1, it means that the block is transmitted, and when it takes 0, it means that the block is not transmitted; after adding up the block vectors transmitted to each user, the total transmission vector w is obtained; each non-zero element of w represents the number of times the block is requested in the current period, and it needs to be transmitted only once;
[0013] Step 3, calculate the average transmission traffic of the system after the line-of-sight guidance, and the degree to which the viewing point position after the guidance deviates from the user's viewing point position before the guidance, and then minimize the weighted sum of the average transmission traffic of the system and the user's viewing point deviation to optimize the line-of-sight guidance position;
[0014] Among them, the objective function for minimizing the weighted sum of the average transmission traffic of the system and the user's viewing point deviation is as follows:
[0015] min J(z)=R(z)+λD(z)
[0016] where 0≤Re(z)≤1, 0≤Img(z)≤1
[0017] R(z) represents the average transmission traffic of the system, D(z) represents the sum of the viewing point deviation distances of all users, λ is the weight, and J(z) is the weighted sum of the average transmission traffic of the system and the user's viewing point deviation; Re(z) and Img(z) are the real part and the imaginary part of z respectively, corresponding to the longitude and latitude coordinates of the guiding position;
[0018] Step 4, the server pushes the optimized line-of-sight guidance position z * to the client and displays it to the user for the user to select.
[0019] The advantages and positive effects of the present invention are as follows: The system and method of the present invention combine the multicast technology to enhance the concentration of the user's viewing field of view, thereby reducing the transmission traffic of the panoramic video. At the same time, the optimized scheme for transmitting the line-of-sight guidance position of the panoramic video adopted by the system and method of the present invention takes into account the user's viewing preferences while comprehensively considering the influence of the line-of-sight guidance on saving transmission traffic and maintaining the user's viewing preferences, and can effectively solve the huge overhead problem of the current panoramic video transmission traffic. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 It is an overall implementation flowchart of the panoramic video transmission method based on line-of-sight guidance of the present invention;
[0021] Figure 2 It is a flowchart for solving the optimized line-of-sight guidance position in the panoramic video transmission method based on line-of-sight guidance of the present invention;
[0022] Figure 3 It is a traffic-user preference graph of the method of the present invention and the existing method under the change of weight λ;
[0023] Figure 4 It is a comparison graph of the traffic savings and user preference deviation in different panoramic video transmissions between the method of the present invention and the existing method. Detailed implementation manners
[0024] The present invention will be further described below in conjunction with embodiments and the accompanying drawings.
[0025] The present invention proposes a panoramic video transmission system based on line-of-sight guidance. Compared with the existing panoramic video transmission system based on slicing, the panoramic video transmission system based on line-of-sight guidance of the present invention adds a line-of-sight guidance point optimization and distribution module in the server and a line-of-sight guidance result presentation module in the client. Its working mode is as follows: The client reports the current viewpoint position to the server; the line-of-sight guidance point optimization and distribution module of the server obtains the viewing characteristics of the user according to the current and historical viewpoint data of the user, and then comprehensively considers the impact of line-of-sight guidance on the transmission traffic and the user's viewing preference, optimizes the line-of-sight guidance point for each user, and distributes the guidance point to the client in real time; after receiving the guidance point, the line-of-sight guidance result presentation module selects an available guidance method in the head-mounted display for presentation, where the guidance method can be designed according to the video content or application scenario. For example, in the form of a player position map in a game scenario, all users' real positions are marked in a small map in the head-mounted display, or the hot spots expected to be guided are marked, so as to attract the user's attention and guide the user's line of sight.
[0026] The line-of-sight guidance point optimization and distribution module in the server establishes a statistical model of the user's viewpoint position after line-of-sight guidance according to the user's current viewpoint, the probability of the user receiving line-of-sight guidance, and the user's line-of-sight guidance habit after receiving the guidance, and calculates the average transmission traffic of the system after line-of-sight guidance, and the degree of deviation of the viewpoint position after guidance from the user's viewpoint position when not guided, and then minimizes the weighted sum of the average transmission traffic of the system and the user's viewpoint deviation to optimize the line-of-sight guidance position. Specifically, the implementation scheme for the line-of-sight guidance point optimization and distribution module to realize the optimization of the line-of-sight guidance position is as described in the first part of the method of the present invention below.
[0027] Specifically, the gaze guidance result presentation module of the client can selectively display on the mini - map in the user's headset:
[0028] (1) The current fixation points of multiple users, guiding the user to move towards the current popular viewing area;
[0029] (2) The fixation points of historical users in the next guidance cycle, guiding the user to move towards the next - second popular area;
[0030] (3) The comprehensively optimized guidance area.
[0031] The first two types of fixation points can be obtained by statistically analyzing user viewpoint data. For the third type, the present invention designs a gaze guidance position optimization scheme considering reducing traffic and maintaining user preferences.
[0032] Consider a panoramic video transmission system. K users watch videos by receiving data transmitted from the same base station. Each spherical frame of the panoramic video can be projected onto a plane. The projected image can be sliced into M blocks and N blocks in terms of longitude and latitude and has normalized longitude - latitude coordinates. In each period available for guidance, first, according to the user's viewpoint p k determine the blocks to be transmitted to user k. In the embodiments of the present invention, the transmitted blocks need to cover the user's 120 - degree field of view. Regarding the viewing area to be transmitted as the union of several blocks, then the vector w(p k ) with dimension N×M can represent the blocks to be transmitted to user k. If the value of each dimension of this vector is 1, then this block is transmitted; if it is 0, it means this block is not transmitted. Then, it is necessary to determine all the transmitted blocks. Using the block transmission for a multicast system, users who request the same block will receive the same multicast data stream, and the remaining blocks requested individually will be transmitted to the corresponding users in a unicast manner to ensure that each block is transmitted at most once. After summing up the vectors w(p k ) of all users during viewing, the transmission vector w is obtained. The value of each element in each dimension of w corresponds to the number of times this block is requested in the current period. Therefore, each non - zero element of the transmission vector w corresponds to a block that needs to be transmitted and only needs to be transmitted once.
[0033] The present invention proposes a panoramic video transmission method based on gaze guidance. First, the client obtains the user's current viewpoint in real - time and sends it to the server. Then, the gaze guidance position is optimized on the server side, and the optimized gaze guidance position is pushed to the client for display on the map in the user's headset. The overall process is as Figure 1 shown. Among them, the first part implements the method for optimizing the gaze guidance position on the server side, specifically as shown in Steps 1 - 2 below.
[0034] Step 1, first calculate the viewpoint position of the user after gaze guidance.
[0035] After gaze guidance, the user's viewpoint position is related to factors such as whether the user accepts the guidance, the user's current viewpoint, the gaze guidance position, the guidance time, and the user's head movement speed. Let p k represent the current viewpoint of user k, and the imaginary and real parts of p k correspond to the longitude and latitude coordinates of the current viewpoint respectively. Let z represent the guidance point or the center of the guidance area. If the user accepts the guidance, the user's viewpoint will move towards the guidance area, and its movement model can be set according to experience or estimated based on the user's viewing behavior; if the user does not accept the guidance, the user's viewpoint will change to whose position can be obtained by means of viewpoint prediction. The guidance point or the center of the guidance area z can be obtained by statistically analyzing the user's viewpoint data, or can be randomly set first and then optimized using the method of the present invention.
[0036] Specifically, a preset model is used to fit the movement of the viewpoint under gaze guidance to determine the viewpoint after the user moves under gaze guidance. For example, a linear movement model is used to fit the head movement process, and it is assumed that within the Δ T time of the continuous guidance, user k will accept the guidance with a probability of φ k . Then, after accepting the guidance, the user will move towards the guidance point or the center of the guidance area z at a head speed of μ k : If the user accepts the guidance, the moved viewpoint satisfies ||.|| represents calculating the modulus of a complex number. Here, it is to normalize the vector z - p k to obtain a unit vector representing the head movement direction. If the user does not accept the guidance, the user's viewpoint will change to the which can be obtained by means of viewpoint prediction. Therefore, the user's viewpoint position after gaze guidance is random. The viewpoint p k and the guidance position z are both represented by complex numbers, with the real part being the longitude coordinate and the imaginary part being the latitude coordinate. For example, (longitude, latitude) = (0.2, 0.3) is represented as 0.2 + 0.3j, where j is the imaginary unit.
[0037] According to the probability of the user accepting gaze guidance and the user's gaze guidance habit after receiving the guidance, a statistical model of the user's viewpoint position after gaze guidance can be established. Specifically, when using the above linear movement model, u k (z) can be used to represent the probability that user k views each block as:
[0038]
[0039] where each dimension u k (z) of u km (z) corresponds to the probability that user k views the m-th block, and the vectors are the blocks transmitted to user k according to the viewpoint respectively.
[0040] Step 2: Calculate the system transmission traffic after line-of-sight guidance and the degree to which the viewpoint position after guidance deviates from the user's viewpoint position before guidance, where the latter reflects the impact of line-of-sight guidance on the user's viewing preference. Based on the viewpoint position of the user after line-of-sight guidance calculated in the previous step and the acceptance probability of the user for the guidance, the user's field of view (FOV) can be calculated, and then the probability \(u\) k (z) can be obtained, where each dimension \(u\) k (z) corresponds to the probability that user \(k\) views the \(m\)-th chunk. To calculate the average transmission traffic, first calculate the probability \(q\) km that none of the \(K\) users views the \(m\)-th chunk: m It is:
[0041]
[0042] Then the probability that at least one of the \(K\) users views the \(m\)-th chunk is \((1 - q\) m ). Since the system uses multicast transmission, no matter how many users view a chunk simultaneously, the base station will only transmit the chunk once. Thus, the average transmission traffic or the average number of transmission chunks \(R(z)\) that the system needs to transmit can be expressed as the sum of the viewing probabilities of each chunk:
[0043]
[0044] If the size of each video chunk is known, the average transmission traffic that the system needs to transmit can be expressed as:
[0045]
[0046] where \(b\) m represents the size of the \(m\)-th chunk.
[0047] The degree to which the viewpoint position after guidance deviates from the user's viewpoint position before guidance can be measured by the Euclidean distance between two points. The distance metric for user \(k\) can be expressed as:
[0048]
[0049] Then the sum of the viewpoint offset distances \(D(z)\) of the \(K\) users can be expressed as:
[0050]
[0051] Finally, solve the following problem to optimize the line-of-sight guidance position to minimize the weighted sum of the system average transmission traffic and the user's viewpoint offset. The objective function is as follows:
[0052] min \(J(z)=R(z)+\lambda D(z)\)
[0053] where \(0\leqslant\mathrm{Re}(z)\leqslant1\)
[0054] and \(0\leqslant\mathrm{Im}(z)\leqslant1\)
[0055] where \(\mathrm{Re}(z)\) and \(\mathrm{Im}(z)\) are the real part and the imaginary part of \(z\), corresponding to the longitude and latitude coordinates of the guiding point
[0056] Pre - set \(\lambda\) as a weight to adjust the importance of traffic conservation and user preference maintenance. After designing a method to solve the objective function, the guiding position can be determined, and the performance evaluation of the line - of - sight guidance in terms of traffic conservation and user preference maintenance can be carried out.
[0057] The non - convexity of the objective function makes the solution difficult: on the one hand, \(R(z)\) is non - convex with respect to \(u\) km (z), on the other hand, \(u\) km (z) is non - convex with respect to \(z\), and it is difficult to obtain the global optimal solution. The traditional gradient - descent method uses the gradient to find the direction in which the objective function decreases fastest, but the condition for obtaining the gradient is that the function is differentiable. However, the above - mentioned objective function has points with extremely large derivatives and does not satisfy the differentiability condition. In addition, the objective function also has points where the derivative is zero, making it difficult to update the gradient. Therefore, this method designs an optimization method that can fit the gradient - descent, and the process is as Figure 2 shown below:
[0058] Step a: Differentially fit the gradient of the objective function. The positions where the objective function drops sharply, that is, the ideal gradient update directions, are random.
[0059] First, randomly generate \(S\) directions. In the embodiment of the present invention, \(S = 8\), and the \(s\) - th direction vector is denoted as where \(\epsilon\) is a relatively small step size. In the embodiment of the present invention, \(\epsilon\) is set to \(0.006\), and the random phase \(\theta\) s can follow a uniform distribution on \((0,2\pi)\). For each direction, calculate the difference \(\eta\) s (z (i) ) as follows:
[0060]
[0061] where \(z\) (i) represents the guiding point of the \(i\) - th iteration.
[0062] Then select the maximum difference value as the fitted gradient, and assume that the corresponding direction vector is s * corresponding to the label of the optimal direction, indicating the direction with a moving step size of \(\epsilon\).
[0063] Step b: Set the jump - out threshold In an embodiment of the present invention, take When the fitted gradient is less than the jump-out threshold , randomly select a direction and jump out with a relatively large step size ∈′, so as to obtain the fitted gradient as:
[0064]
[0065] In an embodiment of the present invention, take ∈′ = 0.2. Φ(z (i) ) represents the fitted gradient.
[0066] Step c, iteratively update the line-of-sight guidance position.
[0067] For the guidance point z (i) in the i-th iteration, the formula for updating the (i + 1)-th guidance point z (i+1) is:
[0068]
[0069] where Proj(·) is a projection function to ensure that the coordinates of the updated guidance point are still within the panoramic image; the projection function can be selected as Proj(z) = min(1, max(0, Re(z))) + j min(1, max(0, Img(z))). σ is a hyperparameter corresponding to the coefficient of the learning rate. ε is a given hyperparameter, and ε = e -8 . m P represents the momentum term of gradient descent, and corrects the current gradient with the historical gradient m′ P :
[0070] m P = β1m′ P +(1 - β1)Φ(z (i) )
[0071] where β1 is used to update the biased first-order moment estimate, and β1 = 0.9 can be taken.
[0072] v P is used to implement the adaptive learning rate, and corrects it with the historical learning rate v′ P :
[0073] v P = β2v′ P +(1 - β2)Φ(z (i) ) 2
[0074] where β2 is used as a damping factor to update the biased second-order moment estimate, and β2 = 0.999 can be taken. The initial values of m P and v P are both 0.
[0075] Preset the maximum number of iterations, and repeat steps a to c until the number of iterations reaches the specified upper limit. Select the position that minimizes the objective function during the iterations as the optimized guiding point position z. * 。
[0076] Through experimental verification, when the weight λ of the objective function > 5, that is, when fully considering the maintenance of user preferences, the optimized guiding point can achieve the same traffic saving effect as the inefficient traversal algorithm, and compared with the traffic without introducing line-of-sight guidance, the ratio can be reduced to less than 0.7.
[0077] The second part of the method of the present invention is to push the optimized line-of-sight guiding point position z * to the client and display it in the mini-map in the user's headset for the user to select. The optimized line-of-sight guiding position of the present invention can achieve the purpose of reducing traffic and maintaining user preferences. At the same time, various line-of-sight guiding position strategies can also be set on the client, such as guiding the user to move to the current popular viewing area according to the current fixation points of multiple users as described above; guiding the user to move to the next second's popular area according to the fixation points of historical users in the next guiding cycle. In this way, different types of line-of-sight guiding position strategies can be set on the client according to the user group in the application scenario to achieve personalized support.
[0078] Compare the system and method of the present invention with the existing heuristic methods through experiments. The comparison results are as Figure 3 and 4 shown. In the figure, the optimization method represents the panoramic video transmission method based on line-of-sight guidance of the present invention; the hot-spot method is the panoramic video transmission method that takes the current viewing hot spot for line-of-sight guidance; the center method is the panoramic video transmission method that takes the center of the picture for line-of-sight guidance; the average method is the panoramic video transmission method that takes the average value of the user's viewpoint coordinates for line-of-sight guidance; the traversal method is the method of obtaining the optimal solution by exhaustive search with each 0.01×0.01 unit on the picture as the guiding point.
[0079] As Figure 3 shown, study the influence of the change of λ in a larger range on the traffic and the preference distance. Change λ, use the optimization method and the traversal method, etc. to find the guiding points respectively, draw the traffic-preference graph, and observe the influence of the change of λ on the traffic and the preference distance. Figure 3 The difference between λ of two adjacent points on the curve in the figure is 5. When λ > 5, the performance of the method of the present invention is almost equal to that of the traversal method, and the curve almost coincides with the traversal method. Further analyze the change trend of the performance of the method of the present invention with λ. From Figure 3It can be seen that the slope of the curve decreases gradually with the increase of λ, which means that: on the one hand, the increase of λ can save the user preference distance within a certain range, but continuing to increase λ will only increase the transmission traffic and cannot exchange for a further reduction in the user preference distance; but on the other hand, in the range of smaller slope of the curve, such as λ>25, reducing λ will cause the corresponding data point to move to the left, indicating that the traffic item has decreased, and at the same time, the data point has not moved up significantly, indicating that the increase in the preference item is not large. In this range, reducing λ can significantly reduce the traffic without excessively increasing the preference distance, that is, using a smaller user preference loss in exchange for a larger traffic saving, while not excessively increasing the user preference distance, providing users with a clearer and smoother viewing experience.
[0080] like Figure 4 As shown in the figure, experiments are conducted on three types of panoramic video transmission. Static videos have the characteristic that the position of the RoI (region of interest) on the screen hardly changes over time, such as interview videos. Dynamic videos have the characteristic that the position of the RoI on the screen changes over time, such as sports videos. The characteristic of exploratory videos is that there is no clear RoI on the screen to attract the user's viewpoint, and the user's viewing tends to explore, such as street view videos. Figure 4 The left figure is a comparison of the traffic ratio without introducing line of sight guidance. It can be seen that the method of the present invention saves more transmission traffic than other methods. The right figure shows the user preference offset on different videos. It can be seen that the method of the present invention has the smallest offset compared with other methods and has more advantages in maintaining user preferences.
Claims
1. A panoramic video transmission system based on line-of-sight guidance, which is improved based on the multicast transmission system of panoramic video chunks, is characterized in that In the panoramic video transmission system based on line-of-sight guidance, a line-of-sight guidance point optimization and distribution module is set in the server, and a line-of-sight guidance result presentation module is set in the client; The client obtains the user's current viewpoint in real time and sends it to the server; Based on the user's current viewpoint, the probability that the user accepts line-of-sight guidance, and the line-of-sight guidance habit after the user receives the guidance, the line-of-sight guidance point optimization and distribution module of the server establishes a statistical model of the user's viewpoint position after line-of-sight guidance, calculates the average transmission traffic of the system after line-of-sight guidance, and the degree to which the viewpoint position after guidance deviates from the user's viewpoint position without guidance. Then, it optimizes the line-of-sight guidance position by minimizing the weighted sum of the average transmission traffic of the system and the user's viewpoint deviation, and sends the optimized line-of-sight guidance position to the client; The line-of-sight guidance result presentation module of the client displays the line-of-sight guidance position obtained from the server to the user for the user to select.
2. The system according to claim 1, wherein The line-of-sight guidance point optimization and distribution module establishes a statistical model of the user's viewpoint position after line-of-sight guidance, including: First, determine the sightline guiding position z, and assume that the current viewpoint of user k is p k , user k has probability φ k Accept the guidance and move towards z. Let the viewpoint after moving be If the user does not accept the guidance, the user's viewpoint will change to the k Prediction Viewpoint Both the viewpoint and the sight guide position are complex numbers, with the real and imaginary parts corresponding to the longitude and latitude coordinates of the position, respectively; Then, according to the probability that the user accepts the gaze guidance and the gaze guidance habit of the user after receiving the guidance, a statistical model of the user's viewpoint position after the gaze guidance is established; let u k (z) represent the probability that user k views each block, as follows: Among them, the vector respectively represents the chunk transmitted to user k according to the viewing point .
3. The system according to claim 1 or 2, characterized in that, The objective function of the line-of-sight guidance point optimization and distribution module for minimizing the weighted sum of the average transmission traffic of the system and the user's viewpoint deviation is as follows: min J(z)=R(z)+λD(z) where 0≤Re(z)≤1, 0≤Img(z)≤1 where z represents the line-of-sight guidance position, R(z) represents the average transmission traffic of the system, D(z) represents the sum of the viewpoint deviation distances of all users, λ is the weight, and J(z) is the weighted sum of the average transmission traffic of the system and the user's viewpoint deviation; Re(z) and Img(z) are the real part and the imaginary part of z, corresponding to the longitude and latitude coordinates of the line-of-sight guidance position.
4. A panoramic video transmission method based on line-of-sight guidance, characterized in that, It includes the following steps: Step 1, the client obtains the user's current viewpoint in real time and sends it to the server; Step 2, the server calculates the user's viewpoint position after the current line-of-sight guidance, including: First, determine the gaze guidance position z. Let the current viewpoint of user k be p k , and user k accepts the guidance with probability φ k and moves towards z. Let the viewpoint after movement be If the guidance is not accepted, the user's viewpoint will change to the predicted viewpoint k based on p Both the viewpoint and the gaze guidance position are complex numbers, and the real and imaginary parts respectively correspond to the longitude and latitude coordinates of the position. Then, based on the probability of the user accepting the gaze guidance and the user's gaze guidance habit after receiving the guidance, establish a statistical model of the user's viewpoint position after the gaze guidance; Among them, the spherical screen of the panoramic video is projected onto a plane and cut into M blocks and N blocks in terms of longitude and latitude. A vector w(p k ) of dimension N×M is used to record the blocks transmitted to user k. The transmitted blocks need to cover the user's field of view. When the value of each dimension of the vector w(p k ) is 1, it means that the block is transmitted, and when the value is 0, it means that the block is not transmitted; the block vectors transmitted to each user are summed to obtain the total transmission vector w; each non-zero element of w indicates the number of times the block is requested in the current period, and it needs to be transmitted only once. Step 3, calculate the average transmission traffic of the system after line-of-sight guidance, and the degree to which the viewpoint position after guidance deviates from the user's viewpoint position without guidance. Then, minimize the weighted sum of the average transmission traffic of the system and the user's viewpoint deviation to optimize the line-of-sight guidance position; Among them, the objective function of minimizing the weighted sum of the average transmission traffic of the system and the user's viewpoint deviation is as follows: min J(z)=R(z)+λD(z) where 0≤Re(z)≤1, 0≤Img(z)≤1 R(z) represents the average transmission traffic of the system, D(z) represents the sum of the viewpoint deviation distances of all users, λ is the weight, and J(z) is the weighted sum of the average transmission traffic of the system and the user's viewpoint deviation; Re(z) and Img(z) are the real part and the imaginary part of z, corresponding to the longitude and latitude coordinates of the line-of-sight guidance position; Step 4, the server pushes the optimized gaze guidance position z * to the client and displays it to the user for the user to select.
5. The method according to claim 4, characterized in that, In the aforementioned step 2, a statistical model of the user's viewpoint position after establishing line-of-sight guidance is established, and u k (z) represents the probability that user k views each block, as follows: Among them, the vector respectively represents the block transmitted to user k according to the viewpoint .
6. The method according to claim 4 or 5, characterized in that In step 2 described above, use a preset model to fit the movement of the viewpoint guided by the line of sight, and determine the viewpoint after the user moves under the guidance 7. The method according to claim 4, wherein In step 3, the average transmission traffic of the system after line-of-sight guidance is calculated as follows: Suppose there are a total of K users, and the probability that none of the K users watches the m-th segment where u km (z) is the probability that user k watches the m-th segment; The probability that the m-th block is viewed by at least one of the users is (1 - q m ); If the sizes of each video block are known, the average transmission traffic R(z) of the system is as follows: Among them, b m represents the size of the m-th block, and MN is the number of blocks.
8. The method according to claim 4, characterized in that In the said step 3, the degree to which the guided viewpoint position deviates from the user's viewpoint position without guidance is measured by the Euclidean distance between the line-of-sight guidance position z and the viewpoint to which the user will change when not accepting guidance. Then, the sum of the viewpoint offset distances of K users 9. The method according to claim 4 or 7 or 8, characterized in that In step 3, an optimization method that can fit gradient descent is used to solve the objective function, including the following: Step a, perform differential fitting on the gradient of the objective function; First, randomly generate S directions, and the s-th direction vector is denoted as where ∈ is the step size, and the random phase θ s obeys a uniform distribution on [0, 2π]; for each direction, calculate the difference η s (z (i) ), as follows: where z (i) represents the line-of-sight guidance position at the i-th iteration; Then select the largest difference value as the fitting gradient, and the corresponding direction is Step b, set the jump-out threshold When the fitted gradient value is less than the jump-out threshold randomly select a direction and jump out with a step size ∈′, where ∈′ is greater than ∈, so as to obtain the fitted gradient as follows: where, Φ(z (i) ) represents the fitting gradient; Step c, update the line-of-sight guidance position; According to the line-of-sight guidance position z in the i-th iteration (i) , update the line-of-sight guidance position z in the (i + 1)-th iteration (i+1) as follows: where Proj(·) is the projection function to ensure that the updated line-of-sight guidance position remains within the panoramic view; σ and ε are given hyperparameters; m P represents the momentum term of gradient descent, and uses the historical gradient m′ P Correction: m P = β1m P + (1 - β1)Φ(z (i) ) where β1 is used to update the biased first-order moment estimate; v P For the adaptive learning rate, the historical learning rate v' is used P Correction: v P = β2v P +(1 - β2)Φ(z (i) ) 2 where β2 is the damping factor, updating the biased second moment estimate; Repeat steps a to c until the maximum number of iterations is reached, and select the position that minimizes the objective function during the iterations as the optimized line-of-sight guidance position z * .
Citation Information
Patent Citations
Method and system for displaying panoramic video
CN104010225A
Information processing device, information processing system, information processing method, and program
CN109845277A