Remote propaganda and distribution system for intelligent recommendation
Through the intelligent recommendation remote publicity and distribution system, combined with the motion, color, texture characteristics and user vision data of the video content, dynamic resource allocation and bit rate resolution optimization are achieved, solving the problems of inflexible resource allocation and fixed configuration dependence in the existing technology, and improving the pertinence and quality of content delivery.
Patent Information
- Application Number
- CN202510334316.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art does not perform detailed analysis of features such as motion, color, texture, etc. when processing video content, which makes it difficult to reflect the importance of the region by resource allocation, and the adjustment of bitrate and resolution depends on preset fixed configurations, and lacks dynamic optimization based on real-time scenes and terminal characteristics.
The remote publicity and distribution system of intelligent recommendation is adopted to calculate the motion speed, color contrast and texture complexity values of the video screen through the content analysis module, generate scene feature index, and generate a region importance matrix based on the region division information. Based on these features, the configuration of code rate and resolution is regulated, and the user's line of sight direction data is collected in real time through the field of view tracking module to optimize resource allocation.
It realizes dynamic adaptability of resource allocation, targeted optimization of bit rate and resolution configuration, improves the targetedness and effectiveness of content delivery, and ensures compatibility and quality consistency between video data streams and terminals during transmission.
Smart Images

Figure CN120151589A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital television, and particularly to a remote publicity distribution system with intelligent recommendation. Background Art
[0002] The technical field of digital television is based on television broadcasting and communication technologies for digital signal transmission and processing, covering the entire technical architecture from content acquisition, compression coding, transmission to reception and playback. Compared with traditional analog television, digital television takes higher transmission efficiency, richer picture quality and sound quality as its core advantages. Through digital encoding technology, video, audio and auxiliary information are converted into digital signals to achieve efficient storage and transmission of multimedia data. The remote publicity distribution system is a solution based on digital television technology, mainly used for accurate distribution of publicity content on a large scale and across multiple terminals. However, in the prior art, when processing video content, features such as motion, color, and texture are not analyzed in detail, resulting in difficulty in reflecting the importance of regions in resource allocation in different scenarios. For the adjustment of bit rate and resolution, the prior art mostly relies on preset fixed configurations and lacks dynamic optimization based on real-time scenarios and terminal characteristics, resulting in the inability to flexibly adjust resource allocation according to scene changes. Therefore, improvements are needed. Summary of the Invention
[0003] The object of the present invention is to solve the deficiencies in the prior art and propose a remote publicity distribution system with intelligent recommendation.
[0004] To achieve the above object, the present invention adopts the following technical solutions: A remote publicity distribution system with intelligent recommendation includes:
[0005] A content analysis module, which calculates the motion speed value, color contrast value, and texture complexity value in the video picture to generate a scene feature index; combines the scene feature index with the region division information under the video frame to generate a region importance matrix;
[0006] A bit rate regulation module, which establishes a mapping relationship between the bit rate value and the video resolution value according to the scene feature index and the region importance matrix, performs regional bit rate allocation on the video data stream to generate a bit rate allocation table; adjusts the numerical set in the bit rate allocation table to generate a bit rate control signal;
[0007] A vision tracking module, which collects the line-of-sight direction data and head movement data in the user's head-mounted device, performs spatial coordinate mapping on the data set to generate the user's vision range, and performs region superposition on the user's vision range and the region importance matrix to generate a vision optimization instruction;
[0008] The publicity and distribution module adjusts the bit rate and resolution parameters in the video data stream based on the bit rate control signal and the field of view optimization instruction, and matches them with the transmission protocol of the digital TV to generate a distribution control instruction.
[0009] Preferably, the steps for obtaining the scene feature index are as follows: calculating the motion speed value, color contrast value, and texture complexity value in the video frame;
[0010] Integrating the motion speed value, color contrast value, and texture complexity value to generate a comprehensive feature data set;
[0011] Based on the comprehensive feature data set, calculate the scene feature index, and the calculation formula is:
[0012]
[0013] Among them, S ci is the scene feature index, D 1 is the motion speed value, D 2 is the color contrast value, D 3 is the texture complexity value.
[0014] Preferably, the steps for obtaining the region importance matrix are as follows: matching the region division information under the video frame with the scene feature index, extracting the number information and pixel density characteristics of all regions in the video frame, analyzing the distribution and density of each region in the video frame, and generating a region pixel density set;
[0015] Based on the region pixel density set, calculate the region weight value, and the calculation formula is:
[0016]
[0017] Among them, W is the region weight value, A p is the pixel density of the current region, QE is the boundary complexity value of the region, QR is the total number of regions, S ci is the scene feature index, A l is the average brightness of the region, L is the boundary length of the region, T is the time stamp value of the video frame, and P is the total number of pixels in the video frame;
[0018] Substitute the region weight value into the relationship of the region numbers to quantitatively assign values to the regions to obtain the region importance matrix.
[0019] Preferably, the steps for obtaining the bit rate allocation table are as follows: performing a region superposition calculation on the region values of the region importance matrix and the scene feature index to form a region characteristic fusion matrix;
[0020] Based on the region characteristic fusion matrix, calculate the region bit rate allocation value, and the calculation formula is:
[0021]
[0022] Among them, R is the regional bitrate allocation value, Q i is the fusion characteristic value of the i-th region in the regional characteristic fusion matrix, F is the video frame rate, H i is the intra-frame complexity value of the i-th region, Z is the scene texture complexity of the current frame, K is the resolution exponent, M is the total bitrate capacity of the video, E i is the intra-frame motion complexity of the i-th region;
[0023] Based on the regional bitrate allocation value, combine the bitrate of each region with the resolution value and allocate it to the corresponding region in the video data stream to generate a bitrate allocation table.
[0024] Preferably, the step of obtaining the bitrate control signal is: based on the bitrate allocation table, check the bitrate value of each region against the corresponding resolution parameter, correct the bitrate value that does not meet the set range in the regional characteristic fusion matrix, and organize the corrected data into a preliminary bitrate adaptation set;
[0025] Based on the preliminary bitrate adaptation set, recalculate the bitrate value of each region item by item, and perform item-by-item correction in combination with the resolution parameter and the current frame rate parameter of the corresponding region in the regional characteristic fusion matrix to generate a regional adjustment bitrate set;
[0026] Based on the regional adjustment bitrate set, gradually screen out the data items whose regional bitrate values exceed the allocation limit range, re-allocate the regional bitrate for the data items that exceed the allocation limit range, and generate a bitrate control signal.
[0027] Preferably, the step of obtaining the user's field of view range is: call the line-of-sight direction data and head movement data recorded by the head-mounted device sensor, classify the line-of-sight direction data according to the time series, and at the same time decompose and vectorize the head movement data, and generate a preliminary direction and movement data set in combination with each time series;
[0028] Based on the preliminary direction and movement data set, convert the two-dimensional angle information of the line-of-sight direction data into three-dimensional coordinate data through three-dimensional space transformation, and at the same time superimpose and correct the displacement and rotation vectors in the head movement data point by point, and perform coordinate mapping on each set of corrected data to generate a three-dimensional space coordinate mapping set;
[0029] Based on the three-dimensional space coordinate mapping set, perform cluster analysis on all mapped three-dimensional coordinate points, divide the user's line-of-sight range according to time segments, and complete multi-dimensional data integration in combination with the line-of-sight direction and movement trend to generate the user's field of view range.
[0030] Preferably, the step of obtaining the field of view optimization instruction is: based on the user field of view and the region importance matrix, one-to-one correspondence is made between the coordinate points of the user field of view and the region numbers of the region importance matrix, each coordinate point is mapped to the corresponding region number and a preliminary region overlay data set is generated;
[0031] Based on the preliminary regional overlay data set, the regional visual field interaction value is calculated, and the calculation formula is:
[0032]
[0033] Among them, X is the regional visual interaction value, R is the set of regional numbers within the user's visual field, Ur is the importance value of the current region, Ir is the overlapping area of the regions within the user's visual field, Tr is the time weight of the current region, and G is the global adjustment parameter;
[0034] Based on the regional visual field interaction value, a visual field optimization instruction is generated.
[0035] Preferably, the step of acquiring the distribution control instruction is: based on the bit rate control signal and the field of view optimization instruction, adjusting the bit rate value of each region in the video data stream frame by frame, and performing adjustment in combination with the resolution parameter of the region, verifying the data adjustment result of each frame according to the standard of the digital television transmission protocol, and generating a protocol matching parameter set;
[0036] Based on the protocol matching parameter set, the regional parameters in the video data stream that have not passed the transmission protocol verification are supplemented and corrected, all parameters that meet the protocol are reorganized and integrated, and a distribution control instruction is generated.
[0037] Compared with the prior art, the advantages and positive effects of the present invention are:
[0038] In the present invention, by refining the motion speed value, color contrast value, and texture complexity value in the video picture, a scene feature index is comprehensively generated, and a regional importance matrix is generated in combination with the regional division information, so as to analyze the characteristics and priorities of different video regions, so that resource allocation has dynamic adaptability, and the configuration of bit rate and resolution is optimized in a targeted manner. Through the real-time collection and spatial mapping of the user's line of sight direction and head movement data, the user's field of view is generated and superimposed with the regional importance matrix, so that regional resource allocation can focus on the user's attention area and improve the pertinence and effectiveness of content delivery. On this basis, the adjusted bit rate and resolution are matched item by item with the digital television transmission protocol, and the distribution parameters are checked and corrected frame by frame to ensure the compatibility and quality consistency of the video data stream with the terminal during the transmission process. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a system flow chart of the present invention. Detailed implementation manners
[0040] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0041] Please refer to Figure 1 , the present invention provides a technical solution: An intelligent recommendation remote publicity distribution system includes:
[0042] A content analysis module that calculates the motion speed value, color contrast value, and texture complexity value in the video picture, generates a scene feature index; combines the scene feature index with the area division information under the video frame to generate an area importance matrix;
[0043] A bitrate regulation module that, according to the scene feature index and the area importance matrix, establishes a mapping relationship between the bitrate value and the video resolution value, performs area bitrate allocation on the video data stream to generate a bitrate allocation table; adjusts the numerical set in the bitrate allocation table to generate a bitrate control signal;
[0044] A vision tracking module that collects the line-of-sight direction data and head movement data in the user's head-mounted device, performs spatial coordinate mapping on the data set to generate the user's vision range, and performs area superposition of the user's vision range and the area importance matrix to generate a vision optimization instruction;
[0045] A publicity distribution module that, based on the bitrate control signal and the vision optimization instruction, adjusts the bitrate and resolution parameters in the video data stream, and matches them with the transmission protocol of the digital TV to generate a distribution control instruction.
[0046] The steps for obtaining the scene feature index are as follows: Calculate the motion speed value, color contrast value, and texture complexity value in the video picture;
[0047] Integrate the motion speed value, color contrast value, and texture complexity value to generate a comprehensive feature data set;
[0048] Based on the comprehensive feature data set, calculate the scene feature index, and the calculation formula is:
[0049]
[0050] Wherein, S ci is the scene feature index, D 1 is the motion speed value, D 2 is the color contrast value, D 3 is the texture complexity value.
[0051] Specifically, calculate the motion speed value, color contrast value, and texture complexity value in the video frame. Combine the previously obtained video source file and its frame-level timing information. Based on the inter-frame difference method, extract the pixel displacement and time interval of the frame changes in several consecutive frames. By recording the pixel coordinates corresponding to each frame and comparing the cumulative displacement values between adjacent frames, obtain the reference value range of the motion speed at the current stage. For example, a pixel displacement of no more than 120 pixels per second is counted as a low speed range, a displacement greater than 120 and no more than 300 pixels is counted as a medium speed range, and a displacement greater than 300 is listed as a high speed range. Here, use frame-by-frame scanning and statistical tools to record all displacement amounts and aggregate them. At the same time, perform a vertical comparison of the position changes after superimposing each frame to determine the speed distribution of each time segment. Subsequently, for the process of obtaining the color contrast value, read the pixel brightness difference degree of the video frame in the RGB three channels. By summing and averaging the brightness differences between each pixel point and its neighboring pixel points, statistically calculate the overall difference degree and map the result to the interval from 0 to 1. Subdivide the difference degree interval exceeding 0.6 to record the high color contrast situation with a value greater than 0.8, and obtain the color contrast statistical array. Then, for the extraction of the texture complexity value, based on the gray-level co-occurrence matrix method, calculate the correlation degree of texture details in different directions from the gray-level distribution relationship of each frame of the image, and then quantitatively score these correlation degrees. For example, define the texture complexity level within the range of 0 to 10, and average the scores of each frame to obtain the current video texture complexity value. When the texture score exceeds 7, it indicates that the texture features are significant, usually meaning that there are a large number of edges or subtle texture changes in the area. After all the data is recorded, arrange them in chronological order according to the frame index and complete the marking synchronously to form a set of speed value sequences, contrast sequences, and texture sequences that are convenient for comparison. Finally, generate the motion speed value, color contrast value, and texture complexity value for the current video segment for subsequent processing.
[0052] Integrate the motion speed value, color contrast value, and texture complexity value. First, through the motion speed value sequence, color contrast statistical array, and texture complexity score set obtained previously, establish a one-to-one correspondence with the video frame index. Select a unified index dimension to ensure that the lengths of all sequences are the same. Then, combine and store the data under the same frame in the established time order to form a comprehensive feature matrix with three columns of fields: speed, contrast, and texture. To reduce the interference of outliers, data points that exceed their respective valid ranges need to be removed. For example, the motion speed value exceeds 1000 pixel displacements or is less than 0 pixel displacements, the color contrast value exceeds 1 or is lower than 0, and the texture complexity value exceeds 10 or is lower than 0. These data points are screened by comparing with the known range thresholds through the inspection program. After that, the remaining data set is divided according to time segments. Combine all the frame data within each second and record it for subsequent feature analysis. If extreme change situations are detected in some time segments, such as the motion speed suddenly increasing to 800 pixel displacements or the texture complexity score exceeding 9, etc., mark them and add remarks for explanation. After forming the complete data structure, the three fields of speed, contrast, and texture corresponding to each time segment are associated by frame, and the original marking information is also recorded in the same comprehensive feature data set table, giving each time period and its corresponding statistical result sequence, thus completing the whole process of merging the three data items of motion speed value, color contrast value, and texture complexity value into a comprehensive feature data set.
[0053] The benefit of the formula is that it comprehensively considers the characteristics of motion speed, color contrast, and texture complexity, and balances the data scales of different dimensions by introducing logarithmic terms and square terms, so as to achieve more-dimensional capture and characterization when describing the scene complexity.
[0054] D 1 The acquisition steps of D are as follows: This parameter represents the motion speed value in the video frame, and it is calculated through the pixel displacement cumulative result and time interval data obtained previously. The specific method is to perform a difference on the pixel coordinate mapping between each frame and the subsequent frame, accumulate the differences and divide by the time interval to obtain the pixel displacement per second. During the actual monitoring process, 3000 frames of data are continuously collected, the pixel displacement situation between each frame is recorded to form a speed sequence, and then the average value and variance of this speed sequence are statistically analyzed to obtain a relatively stable motion speed value. Further compare it with the calibrated actual motion scene to confirm the deviation and ensure correspondence with the objective speed. Finally, select the value of 36.5 as D. 1 The reference input range.
[0055] D 2The acquisition steps are as follows: This parameter represents the color contrast value. The average contrast value and the maximum peak value at the frame level are obtained from the color contrast statistical array obtained previously. The proportion of the contrast between 0 and 1 is calculated using the full-frame pixel brightness difference method, and the highest contrast value of the corresponding frame is included in the reference range. To obtain this parameter, the sum of the brightness differences between each pixel and its neighboring pixels is averaged and then normalized, and the contrast value sequence of all frames is recorded to form a contrast change curve. Based on this curve, the most common peak interval in the central area is selected, and this central interval is compared with other frame segments through software to obtain a weighted average contrast reference value, and a small amount of sampling is linked to correct overly high or low jitter parameters. Finally, the average color contrast value of about 0.72 is obtained as D 2 The input data range of
[0056] D 3 The acquisition steps are as follows: This parameter represents the texture complexity value. The texture score value of each frame is obtained from the texture complexity score set obtained previously. Then, based on the analysis of the gray-level co-occurrence matrix, the occurrence frequencies of different gray-level pairs are quantified, and the distributions of the texture in horizontal, diagonal and other directions are statistically analyzed. Combining edge recognition and noise density to comprehensively evaluate the texture features. In the acquisition process, a video of about 180 seconds is scanned frame by frame. For each frame, several gray-level blocks are extracted and the local contrast, energy and correlation are calculated, and the texture score of the whole frame is obtained by merging. By recording the texture score values of all frames and calculating the average value and median, the texture score items with a sharp local increase are rechecked to determine whether there is a possibility of overly complex picture structure or a large number of fine textures. Finally, the average texture complexity value of 4.1 is obtained as D 3 The reference interval of , and continuous updates are maintained in subsequent monitoring to adapt to different video contents.
[0057] Calculation process:
[0058] First, substitute the numerical values of each parameter:
[0059] D 1 = 36.5, D 2 = 0.72, D 3 = 4.1;
[0060] Perform operations item by item:
[0061] The first step of calculation
[0062]
[0063] The second step of calculation
[0064]
[0065] The third step is to calculate ln(D 3 ):
[0066] ln4.1 ≈ 1.41099;
[0067] Then sum up the three values and multiply by
[0068]
[0069] The result shows that the scene feature index of the current video frame is approximately 2.52623. When the value of S ci is higher than 2.5, it can be regarded as a scene with higher complexity, indicating that there are obvious motion speeds, relatively high color contrasts, and relatively complex texture features in the frame. If the S ci obtained by combining other video segments is used for horizontal comparison, the complexity changes of the frame at different time periods can be further clarified, providing a basis for subsequent remote publicity and distribution strategies.
[0070] The steps to obtain the regional importance matrix are as follows: Match the regional division information under the video frame with the scene feature index, extract the number information and pixel density characteristics of all regions in the video frame, analyze the distribution and density of each region in the video frame, and generate a regional pixel density set;
[0071] Based on the regional pixel density set, calculate the regional weight value. The calculation formula is:
[0072]
[0073] where W is the regional weight value, A p is the pixel density of the current region, QE is the boundary complexity value of the region, QR is the total number of regions, S ci is the scene feature index, A l is the average brightness of the region, L is the boundary length of the region, T is the timestamp value of the video frame, and P is the total number of pixels in the video frame;
[0074] Substitute the regional weight value into the relationship of the regional numbers, and perform quantization assignment on the regions to obtain the regional importance matrix.
[0075] Specifically, based on the comparison between the area division information and the scene feature index under the previously obtained video frames, first read the boundary coordinates of each identified area frame by frame and count the number of pixel points inside the area. At the same time, record the time sequence and picture resolution data corresponding to this frame. Then, perform superposition and comparison on the total number of pixels in each area. Combine the information related to the motion speed, color contrast, and texture features in the previously obtained scene feature index. Compare the sparsity degree of the pixel distribution inside each area with the proportion of the overall pixel number. Perform a ratio operation on the calculated number of pixel points inside the area and the total number of pixel points in the video frame to obtain a pixel proportion sequence to distinguish the density degree of each area in the picture. Summarize the number of pixel points in each area under multiple frames, select the areas with relatively stable distribution as the main positions and list the stable intervals of the pixel proportion. Refer to the total number of pixels when the actual collected video resolution is 1920×1080, which is approximately 2,073,600. Compare the proportion of these areas with the set effective range. For example, when the proportion is lower than 0.01, it is a lower density range; when the proportion is between 0.01 and 0.05, it is a medium density range; when the proportion exceeds 0.05, it is a higher density range. If the pixel proportion of the same numbered area in adjacent frames continuously rises to exceed 0.07, it is recorded as a high density state. Then, record the density values of each area in each frame according to the previously established mapping table. Group the density records of all areas in different frames, and calculate the average value, median, and fluctuation range for the groups. If it is detected that the density values of some areas remain at a relatively high level for a long time, it is determined that their distribution is very concentrated and this situation is noted. While the density values significantly lower than the medium range are marked as sparsely distributed. Through this method of multi-frame superposition and time dimension recording, the pixel density sequences of all areas can be integrated into a unified set and associated with the scene feature index, ultimately forming a comprehensive regional pixel density set.
[0076] The benefit of the formula is to quantitatively describe the importance of the area by integrating multiple dimensions of elements such as the pixel density, boundary complexity, scene feature index, average brightness, area boundary length, and frame time and total number of pixel points inside the area, so that the subsequent process can perform differential processing based on more accurate weight values.
[0077] A pThe acquisition steps are as follows: This parameter represents the pixel density of the current area, and the value can float between 0 and 1. The specific acquisition method is to compare the number of pixel points in the area obtained previously with the total number of pixel points in a single-frame image, and then conduct a summary analysis among frames at different times. By actually detecting multiple segments of videos and randomly selecting several frames per second for pixel statistics, a ratio sequence of the number of pixels in the current area to the total number of pixels in the video frame is obtained for each frame. Subsequently, the average operation is performed on each ratio and data points with abnormal sudden increases and decreases are excluded to form the pixel density value of this area. In addition, the minimum and maximum values of the pixel proportion of the area with the same number are recorded during 300 consecutive frames of monitoring to analyze whether there is abnormal jitter. Finally, A is obtained in one investigation. p = 0.035. This value indicates that the number of pixels in this area accounts for approximately 3.5% of the total pixels in the entire frame.
[0078] The acquisition steps of QE are as follows: This parameter represents the boundary complexity value of the area, which can be quantified by recording the degree of tortuosity of the boundary direction, the number of inflection points, and the undulation changes of each segment. This usually involves contour tracking processing after edge feature detection. The specific method is to calculate the edge curvature distribution in a continuous set of boundary coordinates and regard points with curvature greater than a certain value as inflection points, and then conduct comprehensive quantification based on the number of inflection points, the total length of the boundary, and the curvature change range. The more the boundary deforms, the higher the QE value. By actually surveying multiple target areas at 1080p resolution and statistically analyzing the curvature distribution, referring to more than 30 boundary contour records, the tortuosity score of this area is approximately between 2 and 5. Subsequently, weighted analysis is performed on it, and finally the comprehensive evaluation value of QE is formed. Here, QE = 3.2 is obtained.
[0079] The acquisition steps of QR are as follows: This parameter represents the total number of areas. Usually, it is necessary to determine the total number of all areas during the initial segmentation of the current frame image. The specific method is to divide a certain-scale grid according to the video resolution or perform clustering segmentation based on the image content features, and regard each independent connected component as a single area. The set of numbers of all connected components constitutes the total number of areas. In a single detailed detection, the total number of areas in the same frame may reach more than a dozen or even dozens. By comparing with other frames, if it is found that the number of areas in the same frame basically maintains a certain order of magnitude, it can be regarded as a relatively stable segmentation scheme. Finally, this value is recorded when organizing the mapping relationship. In this example, QR = 12 is obtained by dividing the frame image in the current video scene.
[0080] S ciThe acquisition steps are as follows: This parameter represents the scene feature index, which is derived from the comprehensive data of the previously calculated motion speed value, color contrast value, and texture complexity value. It is a comprehensive value obtained through relevant operations and is used to characterize the overall complexity and dynamic information of the picture. The specific process is to first collect the cumulative motion displacement, brightness difference range, and texture scoring results in hundreds of video frames, and then fuse the three through weighted operations and logarithmic processing to form the scene feature index. The previously obtained S ci is approximately 2.53, indicating that this video segment is at an upper-middle level in terms of speed, color, and texture.
[0081] A l The acquisition steps are as follows: This parameter represents the average brightness of the region. It is necessary to count the brightness values of all pixel points in this region in the current frame image, then sum up the brightness of all pixels and divide by the number of pixels in this region to obtain the average value. To improve accuracy, it is necessary to continuously collect video for several seconds in the actual scene, count the brightness values of the same numbered region in multiple frames, and exclude extreme data that is severely overexposed or underexposed. After adjusting the collected brightness values back to the range of 0 to 255 through metrological calibration, the average brightness value of this region in dozens of frames is finally obtained, and A is obtained through comprehensive calculation. l = 128.
[0082] The acquisition steps of L are as follows: This parameter represents the boundary length of the region. By calculating the distance between adjacent contour points one by one in the previously obtained region contour coordinates and accumulating them to obtain the total length. Specifically, the Euclidean distance formula can be used to accumulate in the pixel coordinates of each boundary segment, and then the lengths of all segments are added together to form the total boundary length. If contour jitter or regional shape changes occur in subsequent multi-frame detections, these changes need to be included in the statistics and the length is recalculated. After comparing multiple shots in the same scene, the average boundary length value of this region is determined, and it is recorded here as L = 260.
[0083] The acquisition steps of T are as follows: This parameter represents the timestamp value of the video frame, usually in seconds or milliseconds, and comes from the timing record of the sequential appearance of each frame in the video stream. The shooting device will attach a corresponding timestamp to each frame during shooting. By reading this timestamp data, the position of the frame in time can be determined. In a video segment with a total duration of 300 seconds, the frames can be numbered approximately at a frame rate of 30 frames per second and compared with the actual shooting time one by one. Finally, the second or millisecond value corresponding to the moment when this frame is located is taken. In the current example, seconds are used as the unit and recorded as T = 45.
[0084] The steps to obtain P are as follows: This parameter represents the total number of pixels in a video frame and needs to be calculated based on the resolution. For a resolution of 1920×1080, multiplying the width 1920 by the height 1080 gives 2,073,600. Through on-site collection, it is confirmed that the video encoding is indeed shot at this resolution. If the resolution changes, the value of this parameter needs to be recalculated. Here, it is recorded as P = 2,073,600.
[0085] Calculation process:
[0086] First, substitute the above parameters:
[0087] A p = 0.035, QE = 3.2, QR = 12, S ci = 2.53, A l = 128, L = 260, T = 45, P = 2,073,600;
[0088] Substitute these values into the parentheses in turn:
[0089]
[0090] Add the two together:
[0091] 0.000530 + 0.01604 = 0.01657;
[0092] Take the absolute value and then take the square root:
[0093]
[0094] This result indicates that the weight of the current area in the picture is approximately 0.1287. When the value is greater than 0 and less than 0.3, it can be regarded as a low importance range, indicating that although the area has certain boundary features, the overall proportion is small and the brightness is not prominent. If the value is higher than 0.5, it indicates that the area importance is relatively obvious. Usually, in subsequent operations, higher bitrates or higher resolutions will be allocated to high-weight areas.
[0095] Based on the relationship between the previously obtained regional weight values and the corresponding regional numbers, first summarize the weight calculation results of each region under all frames, make a horizontal comparison of the fluctuations that the regions with the same number may present at different timestamps, record the weight mean value, fluctuation range, and the time period when the peak value appears for each region, and create a corresponding data comparison table in the order of the numbers. By matching the numbers and weight values one by one, a weight sequence can be listed after each number. If the weight values corresponding to some numbers are located in a relatively high range for a long time, they are marked as numbers that may need to be focused on. If the weight values of some numbers are concentrated in a relatively low range, they are listed as general regions. Measure the percentage of the weight fluctuations under different frames, and mark the specific timestamp and weight value when the fluctuation is significant. Subsequently, conduct a secondary verification of the weight information of all numbers, exclude a small number of registration errors, and proofread the correspondence between the numbers and the actual positions of the images. Finally, make the final number-weight relationship table into a two-dimensional matrix form, with the rows representing the number order and the columns representing the distribution intervals of the weights in other relevant dimensions, and attach this matrix to the time series record to form a unified indexing method, and finally obtain the regional importance matrix.
[0096] The steps to obtain the bitrate allocation table are as follows: perform regional superposition calculation on the regional values of the regional importance matrix and the scene feature index to form a regional characteristic fusion matrix;
[0097] Based on the regional characteristic fusion matrix, calculate the regional bitrate allocation value. The calculation formula is:
[0098]
[0099] Among them, R is the regional bitrate allocation value, and Q i is the fusion characteristic value of the i-th region in the regional characteristic fusion matrix, F is the video frame rate, and H i is the intra-frame complexity value of the i-th region, Z is the scene texture complexity of the current frame, K is the resolution index, M is the total bitrate capacity of the video, and E i is the intra-frame motion complexity of the i-th region;
[0100] Based on the regional bitrate allocation value, combine the bitrate and resolution values of each region and allocate them to the corresponding region in the video data stream to generate the bitrate allocation table.
[0101] Specifically, based on the regional importance matrix and the scene feature index obtained previously, first list all the region numbers and the corresponding weight values one by one, and record the comparison parameters between them and the scene feature index, then scan the weights of each region, retrieve the floating data of these weights under different time series or frame indexes, and compare them with the scene feature index in the scene feature index. Select a reference standard to measure the degree of fusion, add or multiply the weight value of each region with the scene feature index item by item, and multiply the weight value by a coefficient and combine it with the sum of the scene feature indexes to obtain a fusion score that can summarize the comprehensive characteristics of the region in the current time period. When the region number remains stable, accumulate the fusion score in each time segment and observe its change curve. If the fusion score is found to be in the same period, If there is a sharp increase in the time, the dense pixel information and motion trajectory distribution of the area in the picture are checked. According to the motion speed measurement and texture scoring results obtained previously, the recorded scores are compared line by line. If the fusion score exceeds the pre-established high threshold, it is concluded that the area may continue to have high dynamic or high detail features. The number and time data can also be annotated at this time and marked to form the focus of subsequent bit rate allocation. If the fusion score is maintained at a low level in multiple observations, it is included in the ordinary area entry. In this way, multiple information such as motion distribution, pixel density, scene texture and brightness are accumulated and compared frame by frame. Finally, all areas are counted to obtain a fusion score record covering the time series. After denoising the record and correcting the possible jump values, the fusion scores corresponding to each area are summarized in the order of numbering to form a regional characteristic fusion matrix with multi-dimensional feature values.
[0102] The benefit of the formula is that it generates a dynamic and precise bitrate allocation value for each region by integrating multiple key indicators such as the fusion feature value in the regional feature fusion matrix, video frame rate, regional complexity, texture, resolution index, total bitrate capacity, and regional motion complexity.
[0103] Q i The acquisition steps of are as follows: This parameter represents the fusion characteristic value of the ith region in the regional characteristic fusion matrix, which comes from the regional characteristic fusion matrix record obtained previously. By weighted combination of the importance score of the region and the scene characteristic index in different time segments or frame sequences, its high and low changes over time are recorded, and extreme outliers that do not conform to the actual shooting situation are eliminated. Finally, the scores that meet the requirements are averaged under the same region number to obtain a relatively stable value that can reflect the fusion characteristics, which is used as Q i The core input of is, when the target area numbered i=5 is actually recorded once, the detection time is 180 seconds, and a total of 5400 frames are selected for statistics. The obtained fusion feature value sequence is weighted averaged to obtain Q5 ≈0.78, which reflects a relatively high level of fusion characteristics in the current scene for this area.
[0104] The steps to obtain F are as follows: This parameter represents the video frame rate, which is obtained by actually querying the settings of the shooting device or reading the recorded encoding information. Generally, common frame rates in non-high-speed scene shooting are 25fps or 30fps, and can also be configured as 60fps or higher in motion scenes. The larger the frame rate, the more image frames can be processed per unit time. In the current example, 30fps is selected as the normal shooting frame rate, and the frame rate value in the detected video stream is recorded as F = 30.
[0105] H i The steps to obtain H are as follows: This parameter represents the intra-frame complexity value of the i-th area, which is usually quantified by judging the color change degree, edge contour density, and texture overlap degree within a single-frame image. It is necessary to extract the color channel values and texture feature scores of this area in the relevant frames and merge the statistics. When there are multiple frames, the average value is taken to obtain the final intra-frame complexity. If there are large fluctuations in the luminance or color channels during the sampling process, additional records need to be made and marked during the complexity calculation process. In this example, more tortuous edges and texture overlaps are detected in the area numbered i = 5, and after statistics, H 5 = 4.2.
[0106] The steps to obtain Z are as follows: This parameter represents the scene texture complexity of the current frame, which can be obtained from the texture analysis results of the entire frame. After detecting the gray-level co-occurrence matrix and edge features of each frame image in a large range, the texture scores of each local block are summarized to obtain an overall value. If the texture is relatively simple, the score is low; if the edges are rich and the dispersion is high, the value is high. Here, a frame is selected as a representative in the measured video, and after statistics and aggregation, it is confirmed that its texture complexity score is Z = 4.1.
[0107] The steps to obtain K are as follows: This parameter represents the resolution index, a value set to quantify the differences in detail performance at different resolutions. It can be statistically segmented in combination with the actual video encoding resolution used (such as 1920×1080, 1280×720, etc.) and the pixel ratio, and a relatively balanced mapping factor is found through a series of calculations. After comparing various resolutions with sensory clarity and information fidelity, a resolution index that fluctuates within the range of 1.0 to 2.0 is obtained. In this example, the recorded resolution configuration is 1920×1080, and through multiple calculations, K = 1.3 is obtained.
[0108] The steps to obtain M are as follows: This parameter represents the total bitrate capacity of the video, which usually needs to be measured and confirmed according to the bandwidth conditions and encoding scheme. By querying or testing the encoding parameters of the video scene on the shooting device, the theoretical peak bitrate upper limit can be obtained. In this example, the total bitrate under common encoding settings is set to 5000 kbps. After measuring and recording multiple times with a dedicated software, it is found that the actual value fluctuates up and down in the range of 4000 kbps to 4500 kbps. A relatively stable value M = 4500 (unit: kbps) that can support most frame traffic is comprehensively selected.
[0109] E i The steps to obtain E are as follows: This parameter represents the intra-frame motion complexity of the i-th region, which refers to the distribution of motion vectors of adjacent pixel blocks within the region and the inter-frame jump amplitude. It is necessary to compare the motion speed, pixel displacement, etc. of this region before and perform superposition statistics based on the time series. Whenever a large motion vector is detected, the corresponding score is increased. If the motion is basically stable, the score is relatively low. Then, the distribution mean value is used as the quantization value of the intra-frame motion complexity. In this example, the region with the number i = 5 is located in a relatively active range in the scene. After recording the peak values of multiple motion vectors, E 5 = 12.5.
[0110] Calculation process:
[0111] First, substitute the values:
[0112] Q 5 = 0.78, F = 30, H 5 = 4.2, Z = 4.1, K = 1.3, M = 4500, E 5 = 12.5;
[0113] First step, calculate the numerator:
[0114] Q 5 ·F = 0.78 × 30 = 23.4;
[0115] Second step, calculate the denominator part 1:
[0116]
[0117] Divide the numerator by the denominator 1:
[0118]
[0119] Third step, calculate the logarithmic term:
[0120] lnK = ln1.3 ≈ 0.26236;
[0121] Fourth step, calculate the denominator part 2:
[0122] M + E 5 = 4500 + 12.5 = 4512.5;
[0123] Divide the logarithmic term by the denominator 2:
[0124]
[0125] Finally, add the two results and take the absolute value:
[0126] 3.759 + 5.816×10 -5 ≈ 3.75906;
[0127] Therefore:
[0128] R = |3.75906| ≈ 3.75906;
[0129] This result indicates that for the i = 5 area, its bitrate allocation value is approximately 3.75906. The value is within the conventional range of 0 to several tens, reflecting that when the fusion characteristics are high, the frame rate is large, and both the intra-frame motion and texture complexity are obvious, a relatively high bitrate allocation range needs to be arranged for this area in the subsequent resource allocation stage. When R is greater than 3, it can be regarded as a medium to high-level allocation requirement. If it is observed later that R is above 5, further inclination needs to be made in the bitrate resource distribution.
[0130] Based on the recorded area bitrate allocation values obtained previously, first retrieve the values obtained for each area in several time segments in the order of the numbers, and list the corresponding frame rate and picture resolution information. Combining the resolution index defined in the steps and the monitored total bitrate capacity, select the average allocation level for the same area from the previously obtained R values, compare the highest and lowest points in each time segment. If it is found that R is in the medium range in some time segments, it is marked as the normal allocation level. If it is found that R exceeds the predetermined standard, the allocation is increased. If it is found that R is in the lower range, the allocation is decreased. Whenever the detection of a time segment ends, record the bitrate distribution executed when the allocation for this segment is completed, associate it with the video resolution correspondence table, and then map it back to the original area information using the number index. After that, complete this process for all areas, summarize the recorded bitrate and resolution combination data row by row, select the areas with more obvious fluctuations as separate comparisons, and repeat the previous round of allocation process for them to correct inconsistent bitrate values. If a situation where the resolution and bitrate do not match is encountered during the correction, re-determine a reasonable allocation according to the previously defined resolution index and the texture complexity correspondence table for the corresponding scenario. Finally, compile the bitrate allocation completion status of all areas into a set of tables, with the number column corresponding to the area numbers, the bitrate column corresponding to the specific allocation values, the resolution column corresponding to the image resolution parameters used under this value, and indicate the applicable time segment range to generate the bitrate allocation table.
[0131] The steps for obtaining the bitrate control signal are as follows: Based on the bitrate allocation table, check the bitrate value of each region against the corresponding resolution parameter, correct the bitrate values that do not meet the set range in the regional characteristic fusion matrix, and organize the corrected data into a preliminary bitrate adaptation set.
[0132] Based on the preliminary bitrate adaptation set, recalculate the bitrate value of each region item by item, and perform item-by-item correction in combination with the resolution parameter and the current frame rate parameter of the corresponding region in the regional characteristic fusion matrix to generate a regional adjustment bitrate set.
[0133] Based on the regional adjustment bitrate set, gradually screen out the data items whose regional bitrate values exceed the allocated limit range, reallocate the regional bitrate for the data items that exceed the allocated limit range, and generate a bitrate control signal.
[0134] Specifically, based on the previously obtained bitrate allocation table, first read the region numbers and their corresponding bitrate values and resolution parameters, and compare these data item by item with the bitrate range set for this region in the previously obtained regional characteristic fusion matrix. If it is found that the bitrate values of some regions are less than the lowest threshold or greater than the highest threshold marked in the fusion matrix, adjust them according to the previously determined correction criterion, which is jointly determined by experience and subjective evaluation data of the picture quality. For example, in a scene with high requirements for visual fineness, the bitrate is usually allowed to fluctuate between 800 kbps and 2000 kbps, while for regions with simple backgrounds or static pictures, the lower limit of the standard range can be reduced to 300 kbps. If it is found during the investigation that the bitrate of a certain region continuously falls below the minimum value of the standard range in multiple detections, it is marked as an item to be adjusted, and its bitrate is adjusted towards the middle range according to the actual resolution reference value corresponding to the region number. If there is still an obvious deviation after one check, record the deviation amplitude of this time and accumulate it to determine whether the region needs to be further raised to a higher bitrate section or lowered to a lower gear during subsequent integration. To facilitate the program's execution, the bitrate configuration of each region in adjacent frames will be compared coherently. By counting the data of two or three consecutive frames into a timing comparison table to observe its fluctuation pattern, if the fluctuation exceeds the preset threshold, such as more than 0.2 times the bitrate difference, it is listed as an unstable state and then corrected twice. After completing the comparison work for all regions, remap the corrected bitrate values to the corresponding resolutions and arrange them in a preliminary bitrate adaptation set according to the numbers.
[0135] Based on the preliminary bitrate adaptation set obtained earlier, first retrieve the corrected bitrate value of each area in the set and combine it with the resolution parameters and current frame rate parameters under the same number in the regional feature fusion matrix obtained earlier, and compare the texture complexity and motion information presented by each number in the aforementioned time period one by one. If a certain area has high-speed motion features detected in multiple frames, its bitrate is allocated to the upper level to avoid screen ghosting. The one-level allocation here usually refers to the upper limit of the value in the preliminary bitrate adaptation set, and then increase the bitrate by 15% to 20% and record the specific increment, so as to clarify whether a higher bandwidth occupancy is maintained in subsequent time periods. If it is found in the comparison process that the stillness and brightness distribution of the areas with the same number are at a lower complexity For each time period, its bit rate is appropriately reduced to free up more resources for areas with more obvious motion. To facilitate the grading operation, a motion rate comparison table and a brightness threshold interval table are configured for each number. For example, the situation where the average motion speed is less than 30 pixels and the average brightness is stable in the range of 120 to 130 is defined as a low dynamic scene. Therefore, the bit rate usage can be further reduced by 5% to 10% on the basis of the preliminary bit rate adaptation set. By recording the adjustment of all areas in the process in the timing comparison list and saving the final correction value, the comparison can be repeated multiple times until the deviation during the re-inspection converges to the pre-established allowable range. Finally, the bit rates of all areas adjusted frame by frame are aggregated together to form a regional adjustment bit rate set.
[0136] Based on the regional adjustment bit rate set obtained above, the bit rate values of each area are first scanned in each time segment. If it is detected that the bit rate of some numbers rises or falls for multiple frames in a row and exceeds a set allocation limit range, the number is marked separately. This limit range can be graded according to bandwidth load capacity or picture clarity requirements. It is generally divided into three main gears and two transition gears. The specific values are determined by combining the observation of the smoothness of the audience's picture and the statistical data of the equipment in the early test. For example, the interval is set at 500kbps to 1500kbps for the mid-range, 1500kbps to 3000kbps for the high-end, and when the bit rate exceeds 3000kbps or is lower than 500kbps for many times, it is A reallocation operation will be triggered. For the area number that triggers the reallocation, the latest motion and brightness registration records under the same number will be referred to to compare whether the motion rate and brightness average are still within the previous range. If it is found that the motion parameters have been greatly reduced, the bit rate of the area can be switched back to the mid-range and the fluctuation range will be observed in the next detection cycle. If it is found that the motion is still at a high level, the higher bit rate will continue to be used, but the overall bandwidth situation must be considered to decide whether to give up resources to other numbers. After all numbers have been screened through the above cycle, the final stable allocation result that meets the previous threshold limit will be numbered and a record will be regenerated, which contains the bit rate allocation value and resolution correspondence of each number in each time segment to obtain the bit rate control signal.
[0137] The steps for obtaining the user's visual field range are as follows: Call the line-of-sight direction data and head movement data recorded by the head-mounted device sensor, classify the line-of-sight direction data according to the time series, and at the same time decompose and vectorize the head movement data, and generate a preliminary direction and movement data set by combining each time series;
[0138] Based on the preliminary direction and movement data set, convert the two-dimensional angle information of the line-of-sight direction data into three-dimensional coordinate data through three-dimensional space transformation. At the same time, superimpose and correct the displacement and rotation vectors in the head movement data point by point, and perform coordinate mapping on each set of corrected data to generate a three-dimensional space coordinate mapping set;
[0139] Based on the three-dimensional space coordinate mapping set, perform clustering analysis on all mapped three-dimensional coordinate points, divide the user's line-of-sight range according to time segments, and complete multi-dimensional data integration by combining the line-of-sight direction and movement trend to generate the user's visual field range.
[0140] Specifically, call the line-of-sight direction data and head movement data recorded by the head-mounted device sensor, number the line-of-sight direction data according to the time series and list the recording time and the corresponding horizontal and vertical angle values in sequence within each series. At the same time, retrieve the head movement data and disassemble it in the form of displacement and angular velocity. Whenever the head rotation angle or displacement vector at any moment is greater than the preset threshold, an identifier is added to the record. This threshold is obtained by statistically measuring the head rotation rate and angular displacement of multiple users during the normal use of the head-mounted device. Specifically, a large number of time series data can be obtained through multi-person experiments within the first 30 minutes of device wearing, and the maximum and minimum rotation rates and displacement values are extracted. Then, the interval value at two-fifths is taken as the preset threshold during actual application. During this period, the time points also need to be accurately paired to ensure that the line-of-sight direction data and head movement data have a unified time reference. Then, all data information is segmented according to the frame rate, and the line-of-sight angle values and head movement vector values within the same frame or the same time segment are sorted correspondingly. The movement direction and amplitude under each time segment are split and the results are written into the time series list. Then, compare item by item in this list according to the ratio relationship between the line-of-sight angle and head movement. If the line-of-sight angle jumps sharply or the head movement vector generates an abnormal peak, mark it as abnormal at this record and conduct key screening during the next analysis. Finally, reorganize all the line-of-sight angles and head movement information within the normal range into a unified index structure, form an ordered time series data detail, and indicate the angle value, displacement vector, rotation rate, etc. of each record in the field, and generate a preliminary direction and movement data set by combining each time series.
[0141] Based on the preliminary direction and motion data set obtained previously, first, for the horizontal and vertical angle values in the line-of-sight direction data, convert the angles to radians, store the conversion results together with the corresponding time points in an angle list. Then, at each moment, map the horizontal and vertical angles to three-dimensional coordinates according to the three-dimensional space transformation rules. For example, when the horizontal angle is 30 degrees, the vertical angle is 15 degrees, and the range from the center reference direction is no more than 10 degrees, it can be determined that the coordinate vector is biased towards the positive x-axis with a medium-low inclination. If the calculated radian is near π / 2 and the other direction is close to 0, it indicates that the direction is close to directly above. Immediately after each conversion, superimpose the displacement and rotation vectors in the head motion data. If any vector component exceeds the set displacement threshold or rotation threshold at a certain moment, mark it additionally in this record. The setting of these thresholds is obtained from the observation of the actual operation of the head-mounted device and the statistics of the user's motion distribution. For example, set the maximum translation amount not exceeding 0.2 meters and the maximum rotational angular velocity not exceeding 60 degrees per second. If it is detected that the displacement or rotation in a certain sampling is significantly higher than this range, classify this data as high-dynamic recording. Then, according to the distribution of all marked and normal records, fuse the three-dimensional coordinates and the head motion vectors into the same mapping data table, and generate the correction result of the three-dimensional coordinates at the corresponding moment. When performing index matching on each set of corrected data, summarize the results of multiple samplings within the same time segment, and perform fitting confirmation on adjacent sampling points. If some adjacent points are too far apart in three-dimensional space, there may be sampling noise or device drift, and it is necessary to recheck the vector calculation according to the actual motion process. After multiple rounds of correction and screening, finally, form a mapping table arranged in chronological order and containing three-dimensional coordinates and cumulative displacement and rotation information, generating a three-dimensional space coordinate mapping set.
[0142] Based on the three-dimensional space coordinate mapping set obtained previously, first extract the three-dimensional coordinate values at all time points and combine them into a coordinate group, then index them in ascending order according to time. If it is found that some coordinate points appear in clusters multiple times in a short period, record their cluster centers and mark the time periods in the cluster record table. To complete the clustering analysis, start from the Euclidean distance between adjacent coordinate points. If several dense points appear within a certain threshold range, such as within a range of 0.05 meters, they are regarded as one class. If the duration of this class of cluster exceeds 0.5 seconds, it is recorded as a segment where the line of sight is concentrated. By traversing the entire coordinate group, multiple time segments can be detected. Among them, some segments may show continuous movement. At this time, judge whether to merge them into the same movement interval according to the line-of-sight direction and head movement trend obtained previously. If the time difference between two clusters of points is less than a certain threshold, such as 0.3 seconds, and the three-dimensional coordinate position difference does not exceed 0.1 meters, they can be regarded as a longer continuous segment. Divide these continuous segments one by one and incorporate them into the clustering result table. Finally, mark the main line-of-sight area of the user in each time segment in the list, and perform multi-dimensional data integration under the same segment in combination with the corresponding motion vectors. By corresponding the continuous clustering results and rotation vector records, confirm where the line of sight is mainly concentrated in each time period, and then summarize the clustering numbers and coordinate ranges of all segments, list their start and end positions in time, and generate the user's field of view.
[0143] The steps to obtain the field of view optimization instruction are as follows: Based on the user's field of view and the regional importance matrix, map the coordinate points in the user's field of view to the regional numbers in the regional importance matrix one by one, and generate a preliminary regional superposition data set for each coordinate point;
[0144] Based on the preliminary regional superposition data set, calculate the regional field of view interaction value. The calculation formula is:
[0145]
[0146] where X is the regional field of view interaction value, R is the set of regional numbers within the user's field of view, Ur is the importance value of the current region, Ir is the overlapping area of the regions within the user's field of view, Tr is the time weight value of the current region, and G is the global adjustment parameter;
[0147] Generate a field of view optimization instruction based on the regional field of view interaction value.
[0148] Specifically, based on the user's visual field range and the regional importance matrix obtained previously, first extract all the coordinate points within the user's visual field range and compare them one by one with the numbered partitions in the regional importance matrix in chronological order. For each coordinate point, retrieve its actual video frame position. By comparing the coordinate points with the boundary information of each region coordinate by coordinate, if the horizontal and vertical projections of a certain coordinate point and a certain region are both within the boundary range of that region, then map this coordinate point to the corresponding region number. Since the positions of some boundary regions are relatively close, a similarity threshold between the coordinate and the number is established and compared. If the distance from the coordinate point to the center of the region is less than this threshold, it is considered to belong to the same number. This threshold is usually set according to the screen resolution and actual scene measurement. For example, in a 1920×1080 resolution screen, after partitioning, the statistical data shows that a distance less than 50 pixels is considered to belong to the same number. Whenever the matching of the coordinate and the region is completed, record this coordinate point, the corresponding timestamp, and the corresponding region number, and then integrate all the matching results within the same time segment into a temporary comparison table for subsequent retrieval. If a certain coordinate point maintains the same number in multiple consecutive frames, merge the records of these frames to form a superimposed mapping with a longer time span. At the same time, eliminate the records that may have null values or extreme jumps. If it is found that a coordinate point spans multiple region numbers in adjacent frames, mark the number of crossings and the corresponding time period in the record, and then decide whether to remerge these crossing points into the adjacent region numbers according to the specific usage scenario. After the above comparison and merging operations, summarize the corresponding relationships between the coordinates and numbers of all time segments. If there are a large number of coordinate points concentrated in the same number during certain time periods, indicate the degree of aggregation in the superimposed data and record the relevant statistical quantity. Finally, uniformly store and number the pairing results of time sequence, coordinates, and region numbers to obtain the preliminary regional superimposed data set.
[0149] The advantage of the formula is that it jointly participates in the operation through multi-dimensional factors such as the regional importance value, the overlapping area within the user's visual field, the time weight value, and the global adjustment parameter, so as to quantify the user's visual attention in different regions and facilitate more targeted resource allocation in subsequent operations.
[0150] The steps to obtain R are as follows: This parameter represents the set of area numbers within the user's field of view, which contains several specific numbers used to identify different area partitions. The range can be between 1 and dozens of numbers, and the quantity mainly depends on the picture segmentation granularity and the complexity of the actual observation scene. During actual acquisition, all the area numbers covered by the user's line of sight will be counted within the same time period. If a user's line of sight involves multiple numbers during playback, these numbers will be unified into R. To ensure accuracy, it is necessary to perform fine partitioning and number each area during the shooting or rendering stage first, then dynamically record the correspondence between the user's line of sight and the area during the usage stage, and finally match the area numbers frame by frame according to the previously obtained user's field of view and form a deduplicated list of numbers. If 15 area numbers are detected to have an intersection with the user's line of sight position in a 240 - second video clip, then R can be recorded as a discrete set {1, 2, …, 15}.
[0151] The steps to obtain Ur are as follows: This parameter represents the importance value of the current area, which needs to be read from the previously obtained area importance matrix and matched with the index of the corresponding number r. When obtaining it, it is often necessary to comprehensively evaluate indicators such as motion speed, color difference, texture complexity, and brightness characteristics to form a weighted score for this area. This score is then standardized or normalized to obtain the final importance value, and the range is generally between 0 and 1 or 0 and 10, which is obtained through long - term monitoring and integration of multiple measurement methods. If slight fluctuations in the area importance of the same number are found in multiple - frame detections, then an average or median value needs to be used when taking the value to maintain stability. For example, when counting the area numbered 3, after considering multiple factors such as the picture complexity and visual aggregation information, U3 = 0.60 is recorded. This value comes from the quantization process of information such as motion color and texture characteristics in the early stage and has been verified multiple times in multiple test scenarios.
[0152] The steps to obtain Ir are as follows: This parameter represents the overlapping area of the area within the user's field of view, which is used to characterize the covered area or pixel range occupied by this area in the user's actual viewing picture. The specific value is obtained by cross - calculating the boundary of each area number with the coordinate set of the user's line - of - sight projection. For example, for a picture with a resolution of 1920×1080, the pixel rectangle boundaries included in each area can be counted, and then the overlapping part of these pixels with the area where the user's eye movement focus is located can be viewed. After counting the overlapping pixels, they are converted into an area or proportion value. The range usually varies from several hundred to hundreds of thousands of pixels. If the shooting scene is large and the area division is fine, the upper limit of the value may be higher. In practice, the average or peak value of the overlapping area of the same number r can be taken as the core input of Ir by combining multi - frame data. If 5800 pixels that coincide with the user's line of sight are measured for the number 7 in the current time period, then according to the known conversion, I7≈5800.
[0153] The steps to obtain Tr are as follows: This parameter represents the time weight value of the current area, which is obtained by measuring the cumulative duration or critical moments of the area in different time periods. Specifically, the filming device or observation system can record the time periods of each area during the actual playback process during filming or rendering. These time periods are then combined with the moments when the user's line of sight coincides. If the user's line of sight focuses on number r for 10 consecutive seconds, the time weight value of this area may increase. By a certain quantization method, the duration or number of frames is converted into a weight value, and the range is statistically counted and segmented by researchers during scene acquisition. For example, in most scenes, the cumulative appearance duration within 1 to 5 seconds is recorded as a lower interval, 5 to 15 seconds as a medium interval, and more than 15 seconds as a higher interval. Different weight scores can be assigned to each interval. If area number 4 has been monitored and it is statistically found that the user's concentrated line of sight has appeared for 12 seconds within a 30-second video clip, then T4 = 4.5 can be sorted out, and this value is calculated by the reference interval method.
[0154] The steps to obtain G are as follows: This parameter represents the global adjustment parameter, which is mainly used to balance the deviation in the visual field interaction calculation of different areas. In order to obtain a global value applicable to multiple scenarios, a large number of tests are often required in combination with the actual picture complexity and the observer's visual sensitivity. The methods include continuously running the device for more than 200 hours, collecting the distribution of the user's focus points in the picture under different shooting scenarios on a large scale, forming a preliminary global parameter sequence through weighted measurement of multiple indicators of the user's visual preference and picture characteristics, and then obtaining a relatively stable parameter value after extreme value removal and weighted average operations on this sequence. The range is usually between 1 and 10, and appropriate upper and lower corrections can be made under different scenarios. For example, after multiple statistics, G = 2.5 is taken for indoor high-dynamic scenarios, and G = 4 is taken for outdoor large-range shots. In this example, G = 2.5 is recorded.
[0155] Calculation process:
[0156] Here, R is regarded as a discrete set {1, 2, 3} for simplified calculation, and it is set that:
[0157] U1 = 0.45, U2 = 0.82, U3 = 0.60;
[0158] I1 = 4000, I2 = 9800, I3 = 3000;
[0159] T1 = 2.2, T2 = 5.3, T3 = 9.1, G = 2.5;
[0160] Let the integral ∫ R Under discrete conditions, it is converted into a summation form, then:
[0161]
[0162] Step 1, calculate
[0163]
[0164] Step 2, calculate the numerator and denominator for each r respectively:
[0165] For r = 1:
[0166] U1·I1 = 0.45 × 4000 = 1800;
[0167]
[0168]
[0169] For r = 2:
[0170] U2·I2 = 0.82 × 9800 = 8036;
[0171]
[0172]
[0173] For r = 3:
[0174] U3·I3 = 0.60 × 3000 = 1800;
[0175]
[0176]
[0177] Step 3, accumulate the results:
[0178] X = 476.0 + 1167.67 + 168.51 ≈ 1812.18;
[0179] This result indicates that when the three regions have different importance values, overlapping areas, and time weights respectively, the total regional visual field interaction value is approximately 1812.18. A value greater than 1000 indicates that the overall attention to the regions within the user's visual field is relatively high. If more complexity is added in subsequent scenarios or the overlapping area and time weight of a certain region are further increased, the calculation result will also increase, which is used to distinguish high-attention regions from ordinary regions in subsequent processes.
[0180] Based on the previously obtained regional visual interaction values, first retrieve the importance values of each regional number and the recorded user viewing time within the corresponding time segment, and horizontally compare these data with the regional visual interaction values one by one. By observing the magnitude of the regional visual interaction values, identify which regional numbers have a relatively high proportion in the user's current field of view, and query the associated motion or texture information for that number. If the value continuously increases and exceeds the pre-determined high threshold range during several monitoring processes, it is classified into the high-concern regional numbers. If it continuously remains in a lower range compared to other numbers during multiple comparisons, it is marked as a low-concern number, and its numerical fluctuations are continuously tracked in subsequent cycles. Then, combine the current video frame sequence and subsequent transmission requirements to classify and summarize these high-concern and low-concern numbers. By recording the temporal trend of the interaction values of each number in different time segments, if there is a significant change in the cumulative interaction value of a certain number in the recent few segments, list the specific change range and make remarks. After screening and sorting the interaction values of all numbers, summarize the recent attention level classification of each number, and then link it with the picture resolution and bitrate allocation information to further formulate a specific allocation plan. Finally, statistically analyze and number the mapping of the number and the interaction value combined with the user's line-of-sight duration information to generate a visual optimization instruction.
[0181] The steps to obtain the distribution control instruction are as follows: Based on the bitrate control signal and the visual optimization instruction, adjust the bitrate values of each region in the video data stream frame by frame, and at the same time perform adjustments in combination with the resolution parameters of the region. Verify the data adjustment results of each frame according to the standard of the digital television transmission protocol to generate a set of protocol matching parameters.
[0182] Based on the set of protocol matching parameters, supplement and correct the region parameters in the video data stream that do not pass the transmission protocol verification, and reorganize and integrate all the parameters that meet the protocol to generate a distribution control instruction.
[0183] Specifically, based on the bitrate control signal and the field of view optimization instruction obtained previously, first read the bitrate values of each region in the video data stream frame by frame and pair them with the previously recorded resolution parameters. By listing the bitrate values and resolution values corresponding to the region numbers in each time segment, and then comparing with the grading ranges formulated in the previous stage. For example, the bitrate is pre-divided into low grade from 0 kbps to 500 kbps, medium grade from 500 kbps to 2000 kbps, and higher grade from 2000 kbps to 5000 kbps, and corresponding comparison criteria are set for different resolution parameters. If the bitrate jump amplitude of a certain region between adjacent frames is greater than the predetermined standard, then mark the region number additionally and record the current jump value. These standards are usually set in combination with the complexity of the video scene and the bandwidth margin. For example, select a bitrate difference of 0.3 times or 0.5 times as the jump threshold under different test conditions, and obtain it through playing and measuring a number of video samples on the device for 120 hours. Subsequently, after completing the combination of bitrate and resolution for each frame, compare the corresponding region records in the data stream with the standard format of the digital television transmission protocol. The information to be verified here includes whether the actual bitrate exceeds the bandwidth limit, whether the resolution parameters meet the frame resolution consistency requirements, etc. If it is found that a certain region exceeds the specified transmission upper limit during the comparison, list the number and its bitrate at the verification record, and at the same time take a snapshot registration of the marked bitrate difference and resolution values, so that the adjustment link in the next time period can be adjusted with reference to this snapshot. If the same number frequently exceeds the limit after multiple verifications, then include this number in the unstable list. If some numbers are continuously lower than the medium bitrate in multiple frames, they will be classified into the low bitrate list. After all frames are verified, mark these unstable or low bitrate numbers and organize them into a classification index. Finally, summarize all the verified region parameters and their timing positions to obtain the complete comparison result and the corresponding status mark, forming a set of protocol matching parameters.
[0184] Based on the protocol matching parameter set obtained previously, first list the region numbers that failed the transmission protocol check one by one and query the bitrate and resolution snapshots recorded in the previous steps. For values that exceed the limit or are too low, compare the intra-frame pixel complexity and motion rate frame by frame. If the same excessive value or low bitrate value continuously appears within an approximate time segment, mark the number as a key object for correction, and reconfigure its parameters by combining the motion speed and texture score distribution in the previous region feature fusion matrix. In this configuration, two-way adjustment will be made with reference to the mapping relationship between resolution and bitrate. For example, adjust the resolution originally configured in the high-end range down by one level or correct the low-end bitrate upward by a range of 200 kbps to 600 kbps. These adjustment amplitudes are derived from empirical data on inter-frame fluctuations and comparisons of subjective image sharpness evaluations by observers. If the situation of non-matching with the protocol still occurs after correction, continue to record the correction amplitude each time, and track the historical adjustment trajectory of this number at different time periods in the form of a cumulative vector. After completing the correction of all numbers that failed the check, integrate the new bitrate and resolution values once. Renumber and file the numbers that already conform to the protocol in the integration result according to the region order, and map them to the complete time series table for final confirmation. If a number still fails to meet the rules in the last check, list it separately in the troubleshooting list for further in-depth adjustment later. Finally, form a unified allocation interval index for all parameters that conform to the protocol and match it to the frame sequence of the video data stream to generate distribution control instructions.
Claims
1. An intelligent recommendation remote publicity distribution system, characterized in that: The system comprises: The content analysis module calculates the motion speed value, color contrast value, and texture complexity value in the video image to generate a scene feature index; combines the scene feature index with the region division information under the video frame to generate a region importance matrix; A bit rate control module, which establishes a mapping relationship between bit rate values and video resolution values according to the scene feature index and the regional importance matrix, distributes bit rates to the video data stream by region, and generates a bit rate allocation table; and adjusts the value set in the bit rate allocation table to generate a bit rate control signal; The field of view tracking module collects the gaze direction data and head movement data from the user's head-mounted device, maps the data set to spatial coordinates, generates the user's field of view, and superimposes the user's field of view with the regional importance matrix to generate field of view optimization instructions; The publicity distribution module adjusts the bit rate and resolution parameters in the video data stream based on the bit rate control signal and the field of view optimization instruction, matches them with the transmission protocol of the digital television, and generates a distribution control instruction.
2. The remote publicity distribution system of intelligent recommendation according to claim 1, characterized in that: The steps of obtaining the scene feature index are as follows: Calculate the motion speed value, color contrast value and texture complexity value in the video picture; Integrating the motion speed value, color contrast value and texture complexity value to generate a comprehensive feature data set; Based on the comprehensive feature data set, the scene feature index is calculated using the following formula: Among them, S ci is the scene feature index, D1 is the motion speed value, D2 is the color contrast value, and D3 is the texture complexity value.
3. The remote publicity distribution system of intelligent recommendation according to claim 1, characterized in that: The steps for obtaining the regional importance matrix are: Matching the region division information under the video frame with the scene feature index, extracting the number information and pixel density characteristics of all regions in the video frame, analyzing the distribution and density of each region in the video picture, and generating a regional pixel density set; Based on the regional pixel density set, the regional weight value is calculated using the following formula: Among them, W is the regional weight value, A p is the pixel density of the current area, QE is the boundary complexity value of the area, QR is the total number of areas, S ci is the scene feature index, A l is the average brightness of the region, L is the boundary length of the region, T is the timestamp value of the video frame, and P is the total number of pixels in the video frame; Substitute the region weight value into the relationship of the region number, perform quantitative assignment on the region, and obtain the region importance matrix.
4. The remote publicity distribution system of intelligent recommendation according to claim 1, characterized in that: The steps of obtaining the code rate allocation table are as follows: Performing regional superposition calculation on each regional value of the regional importance matrix and the scene feature index to form a regional characteristic fusion matrix; Based on the regional characteristic fusion matrix, the regional bit rate allocation value is calculated, and the calculation formula is: Among them, R is the regional code rate allocation value, Q i is the fusion characteristic value of the i-th region in the regional characteristic fusion matrix, F is the video frame rate, H i is the intra-frame complexity value of the i-th region, Z is the scene texture complexity of the current frame, K is the resolution index, M is the total bit rate capacity of the video, and E i is the intra-frame motion complexity of the i-th region; Based on the regional bit rate allocation value, the bit rate of each region is combined with the resolution value and allocated to the corresponding region in the video data stream to generate a bit rate allocation table.
5. The remote publicity distribution system of intelligent recommendation according to claim 1, characterized in that: The steps of obtaining the bit rate control signal are: Based on the bit rate allocation table, the bit rate value of each region is checked against the corresponding resolution parameter, the bit rate value that does not conform to the set range in the regional characteristic fusion matrix is corrected, and the corrected data is organized into a preliminary bit rate adaptation set; Based on the preliminary bitrate adaptation set, recalculate the bitrate value of each region item by item, combine the resolution parameter and current frame rate parameter of the corresponding region in the regional characteristic fusion matrix, perform correction item by item, and generate a regional adjustment bitrate set; Based on the regional adjustment rate set, data items whose regional rate values exceed the allocation restriction range are gradually screened, regional rate is reallocated to the data items exceeding the allocation restriction range, and a rate control signal is generated.
6. The remote publicity distribution system of intelligent recommendation according to claim 1, characterized in that: The steps for obtaining the user's visual field are as follows: Call the gaze direction data and head movement data recorded by the head-mounted device sensor, classify the gaze direction data by time series, decompose and vectorize the head movement data, and combine each time series to generate a preliminary direction and movement data set; Based on the preliminary direction and motion data set, the two-dimensional angle information of the sight direction data is converted into three-dimensional coordinate data through three-dimensional space transformation, and the displacement and rotation vectors in the head motion data are superimposed and corrected point by point, and each set of corrected data is coordinate mapped to generate a three-dimensional space coordinate mapping set; Based on the three-dimensional space coordinate mapping set, cluster analysis is performed on all mapped three-dimensional coordinate points, the user's visual range is divided according to time segments, and multi-dimensional data integration is completed in combination with the visual direction and movement trend to generate the user's visual range.
7. The remote publicity distribution system of intelligent recommendation according to claim 1, characterized in that: The steps for obtaining the field of view optimization instruction are as follows: Based on the user's visual field and the regional importance matrix, coordinate points of the user's visual field are matched one by one with the regional numbers of the regional importance matrix, each coordinate point is mapped to a corresponding regional number and a preliminary regional overlay data set is generated; Based on the preliminary regional overlay data set, the regional visual field interaction value is calculated, and the calculation formula is: Among them, X is the regional visual interaction value, R is the set of regional numbers within the user's visual range, U(r) is the importance value of the current region, I(r) is the overlapping area of the region within the user's visual range, T(r) is the time weight of the current region, and G is the global adjustment parameter; Based on the regional visual field interaction value, a visual field optimization instruction is generated.
8. The remote advertising and distribution system for intelligent recommendation according to claim 1, characterized in that: The steps of obtaining the distribution control instruction are as follows: Based on the bit rate control signal and the field of view optimization instruction, the bit rate value of each region in the video data stream is adjusted frame by frame, and the adjustment is performed in combination with the resolution parameter of the region, and the data adjustment result of each frame is verified according to the standard of the digital television transmission protocol to generate a protocol matching parameter set; Based on the protocol matching parameter set, the regional parameters in the video data stream that have not passed the transmission protocol verification are supplemented and corrected, all parameters that meet the protocol are reorganized and integrated, and a distribution control instruction is generated.
Citation Information
Patent Citations
Method for extracting regions of interest in real-time video communication
CN104079934A
Multi-resolution virtual-reality equipment screen content encoding algorithm by using eye tracking data
CN107770561A
Method, device, medium and program product for encoding video information
CN116248887A
Preprocessing immersive video
CN117044210A
Scalable audio and video coding method and system
CN118488245A
Cited By
Ultrahigh-definition video 5G aggregation transmission method
CN120750835A
Game texture resource optimization processing method and device, electronic equipment and storage medium
CN121550687A