Visual information transmission deterministic network construction method and system
By building a video transmission network and dynamically adjusting the resource allocation ratio based on real-time feedback video fluctuation characteristics and prediction characteristics, the problem of low cost utilization in the deterministic network of visual information transmission is solved, and efficient and flexible resource allocation and low-latency transmission are achieved.
Patent Information
- Application Number
- CN202510090127.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-13
AI Technical Summary
In visual information transmission deterministic networks, how to provide a deterministic network with low latency at lower costs and improve cost utilization.
By obtaining the transmission performance parameters of the video acquisition end and the video receiver, a video transmission network is built, and the resource allocation ratio of each video acquisition end is determined based on the video transmission network. The video acquisition terminal feedbacks the video fluctuation characteristics and video prediction characteristics in real time, adjusts the resource allocation ratio, and dynamically adjusts resource allocation.
It realizes efficient and flexible resource allocation in practical application scenarios, and improves the cost utilization rate and transmission efficiency of the video transmission network.
Smart Images

Figure CN119996729A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video transmission, and in particular to a method and system for constructing a deterministic network for visual information transmission. Background Art
[0002] Deterministic Network for Visual Information Transmission is a network design method that aims to provide efficient, stable and low-latency transmission for visual data (such as images, videos, three-dimensional scenes, etc.). It is usually suitable for application scenarios that require guaranteed transmission quality, such as video surveillance, autonomous driving, virtual reality, augmented reality, etc.
[0003] In layman's terms, a deterministic network for visual information transmission is a low-latency transmission of video data. When resources are sufficient, it is only necessary to improve the transmission performance. However, in actual applications, each application scenario will have a cost issue. How to improve cost utilization and provide a low-latency deterministic network at a lower cost is the technical problem that the technical solution of the present invention wants to solve. Summary of the invention
[0004] The purpose of the present invention is to provide a method and system for constructing a deterministic network for visual information transmission to solve the problems raised in the above-mentioned background technology.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A method and system for constructing a deterministic network for visual information transmission, the method comprising:
[0007] Acquire the transmission performance parameters of the video acquisition end and the transmission performance parameters of the video receiving end, and build a video transmission network according to the transmission performance parameters of the video acquisition end, the transmission performance parameters of the video receiving end and the distance between the video acquisition end and the video receiving end;
[0008] Determine the resource allocation ratio of each video acquisition terminal based on the video transmission network;
[0009] Receive video fluctuation characteristics and video prediction characteristics fed back in real time by the video acquisition end, and adjust the resource allocation ratio according to the video fluctuation characteristics and video prediction characteristics;
[0010] Allocating transmission resources based on the adjusted resource allocation ratio;
[0011] Among them, the video fluctuation feature and the video prediction feature are both independently extracted by the video acquisition end, the video fluctuation feature is used to characterize the degree of change of each image in the video over time, and the video prediction feature is used to characterize the prediction accuracy of each image in the video; the resource allocation ratio is directly proportional to the video fluctuation feature and inversely proportional to the video prediction feature.
[0012] As a further solution of the present invention: the step of obtaining the transmission performance parameters of the video acquisition end and the transmission performance parameters of the video receiving end, and building a video transmission network according to the transmission performance parameters of the video acquisition end, the transmission performance parameters of the video receiving end and the distance between the video acquisition end and the video receiving end includes:
[0013] Querying the transmission performance parameters of the video acquisition end and the transmission performance parameters of the video receiving end in the port filing data; the transmission performance parameters adopt at least one of bandwidth, delay and throughput;
[0014] Input the transmission performance parameters into a preset conversion model to obtain a virtual distance;
[0015] Get the locations of all video acquisition terminals and all video receiving terminals, and create nodes;
[0016] A directed connection line between nodes is determined based on the virtual distance; the direction of the directed connection line is from the video acquisition end to the video receiving end, and the value of the directed connection line is the virtual distance.
[0017] As a further solution of the present invention: the step of determining the resource allocation ratio of each video acquisition terminal based on the video transmission network includes:
[0018] Count the values of each directed connection line and calculate the total value;
[0019] Calculate the ratio of the value of each directed link to the total value;
[0020] For any video acquisition end, all the proportions corresponding to the video acquisition end are counted to obtain the final resource allocation ratio.
[0021] As a further solution of the present invention: the process of extracting video fluctuation features and video prediction features at the video acquisition end includes the following steps:
[0022] Convert the acquired video into an image set sequence and an audio sequence;
[0023] Identifying the audio sequence, determining an initial separation point, and segmenting the image set sequence based on the initial separation point to obtain an image subset;
[0024] Perform image comparison on each image subset to determine the final separation point;
[0025] Determine the video fluctuation characteristics according to the distribution of the final separation points and the corresponding image comparison results;
[0026] Adjacent separation points to be analyzed are selected from the final separation points, and video prediction features are determined based on images within the adjacent separation points.
[0027] As a further solution of the present invention: the steps of identifying the audio sequence, determining the initial separation point, and segmenting the image set sequence based on the initial separation point to obtain the image subset include:
[0028] Segment the audio sequence according to the audio amplitude to obtain subsequences;
[0029] Input each subsequence into the preset audio recognition model and output the recognized text;
[0030] Merging the subsequences based on the recognized text; wherein, when the recognized text has a dialogue feature, executing the merging process; when the recognized texts belong to the same preset audio media, executing the merging process;
[0031] The first and last moments of each merged subsequence are used as initial separation points;
[0032] The image set sequence is segmented based on the initial separation point to obtain image subsets.
[0033] As a further solution of the present invention: the step of performing image comparison on each image subset to determine the final separation point comprises:
[0034] For each image subset, the images are read sequentially based on the time sequence and the images are grayed out;
[0035] Read the gray value of each pixel in the image according to a preset pixel reading order, and perform functionization on the image;
[0036] Perform Fourier transform on the functionalized image to obtain a frequency domain image;
[0037] For frequency domain images at adjacent moments, select a region of a preset size in the frequency domain image, compare the selected region, and calculate the image similarity;
[0038] Select images whose image similarity is less than a preset similarity threshold, and use the corresponding time point as the final separation point;
[0039] In the process of selecting a region of a preset size in the frequency domain image, the area of the region is inversely proportional to the CPU occupancy rate of the video acquisition end.
[0040] As a further solution of the present invention: the step of determining the video fluctuation characteristics according to the distribution of the final separation points and the corresponding image comparison results includes:
[0041] Calculate the distance between each final separation point and the adjacent final separation point and determine the correction factor;
[0042] Read the image similarity corresponding to each final separation point;
[0043] Determine the fluctuation value of each final separation point according to the correction coefficient and the image similarity;
[0044] The fluctuation values of all final separation points are counted as the video fluctuation feature.
[0045] As a further solution of the present invention: the calculation process of the video fluctuation feature is:
[0046] Where B is the video fluctuation feature, B i is the fluctuation value of the ith final separation point, N is the total number of final separation points, α and β are the preset correction coefficients, S i is the image similarity at the i-th final separation point, min{d(i, i-1), d(i, i+1)) represents the minimum distance, which is the minimum value of the time difference between the i-th final separation point and the i-1-th final separation point and the time difference between the i+1-th final separation point and the i-th final separation point; when i is 1 and N, there is only one adjacent final separation point, and the time difference is directly read as the minimum value.
[0047] As a further solution of the present invention: the step of selecting adjacent separation points to be analyzed from the final separation points and determining the video prediction features according to the images in the adjacent separation points comprises:
[0048] Select two adjacent separation points in the final separation points in turn as adjacent separation points to be analyzed;
[0049] Select the image corresponding to the midpoint of adjacent separation points as the reference image;
[0050] Calculating the similarity between each image in adjacent separation points and the reference image, and calculating the image prediction accuracy based on the similarity;
[0051] The image prediction accuracy corresponding to all adjacent separation points is counted and the video prediction features are calculated.
[0052] As a further solution of the present invention: the process of determining the image prediction accuracy is:
[0053] Where, T z is the image prediction accuracy, S j is the similarity between the jth image in the adjacent separation point and the reference image, and M is the total number of images in the adjacent separation point;
[0054] The calculation process of the video prediction feature is:
[0055] Where, T v is the video prediction feature, T zn Represents the image prediction accuracy of the nth pair of adjacent separation points.
[0056] As a further solution of the present invention: the process of adjusting the resource allocation ratio according to the video fluctuation characteristics and the video prediction characteristics includes:
[0057] Wherein, F′ is the resource allocation ratio after adjustment, c1 and c2 are preset constants, and F is the resource allocation ratio before adjustment.
[0058] The technical solution of the present invention also provides a system for constructing a deterministic network for visual information transmission, the system comprising:
[0059] A network construction module is used to obtain the transmission performance parameters of the video acquisition end and the transmission performance parameters of the video receiving end, and to build a video transmission network according to the transmission performance parameters of the video acquisition end, the transmission performance parameters of the video receiving end and the distance between the video acquisition end and the video receiving end;
[0060] An allocation ratio determination module, used to determine the resource allocation ratio of each video acquisition terminal based on the video transmission network;
[0061] An allocation ratio adjustment module is used to receive video fluctuation characteristics and video prediction characteristics fed back in real time by a video acquisition end, and adjust the resource allocation ratio according to the video fluctuation characteristics and video prediction characteristics;
[0062] An allocation ratio application module, used for allocating transmission resources based on the adjusted resource allocation ratio;
[0063] Among them, the video fluctuation feature and the video prediction feature are both independently extracted by the video acquisition end, the video fluctuation feature is used to characterize the degree of change of each image in the video over time, and the video prediction feature is used to characterize the prediction accuracy of each image in the video; the resource allocation ratio is directly proportional to the video fluctuation feature and inversely proportional to the video prediction feature.
[0064] As a further solution of the present invention: the network construction module includes:
[0065] A parameter query unit, used to query the transmission performance parameters of the video acquisition end and the transmission performance parameters of the video receiving end in the port filing data; the transmission performance parameters adopt at least one of bandwidth, delay and throughput;
[0066] A virtual distance generating unit, used for inputting the transmission performance parameter into a preset conversion model to obtain a virtual distance;
[0067] A node creation unit, used to obtain the positions of all video acquisition terminals and all video receiving terminals, and create nodes;
[0068] The node connection unit is used to determine the directed connection line between nodes based on the virtual distance; the direction of the directed connection line is from the video acquisition end to the video receiving end, and the value of the directed connection line is the virtual distance.
[0069] As a further solution of the present invention: the allocation ratio determination module includes:
[0070] A numerical sum calculation unit is used to count the values of each directed connection line and calculate the sum of the values;
[0071] A ratio calculation unit, used to calculate the ratio of the value of each directed link to the total value;
[0072] The ratio statistics unit is used to count all the ratios corresponding to any video acquisition end to obtain the final resource allocation ratio.
[0073] Compared with the prior art, the present invention has the following beneficial effects:
[0074] The present invention performs performance analysis on the video acquisition end to determine the resource allocation ratio of each video acquisition end. On this basis, the video acquisition end analyzes the acquired video to determine the fluctuation characteristics and predictability of the video, and adjusts the resource allocation ratio in real time according to the fluctuation characteristics and predictability, thereby providing a dynamic resource allocation solution that is highly consistent with the actual situation and highly flexible. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention.
[0076] Figure 1 The overall flow chart of the method for constructing a deterministic network for visual information transmission is shown.
[0077] Figure 2 A structural diagram of a deterministic network construction system for visual information transmission is shown. DETAILED DESCRIPTION
[0078] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0079] Embodiment 1:
[0080] Figure 1 The overall flow chart of the method and system for constructing a deterministic network for visual information transmission is as follows. In an embodiment of the present invention, a method for constructing a deterministic network for visual information transmission includes:
[0081] Step S100: acquiring transmission performance parameters of the video acquisition end and the video receiving end, and building a video transmission network according to the transmission performance parameters of the video acquisition end, the transmission performance parameters of the video receiving end, and the distance between the video acquisition end and the video receiving end;
[0082] Step S200: determining the resource allocation ratio of each video acquisition terminal based on the video transmission network;
[0083] Step S300: receiving video fluctuation characteristics and video prediction characteristics fed back in real time by the video acquisition end, and adjusting the resource allocation ratio according to the video fluctuation characteristics and video prediction characteristics;
[0084] Step S400: Allocate transmission resources based on the adjusted resource allocation ratio.
[0085] The technical solution of the present invention is applied to a local area network, such as a workshop or power scene equipped with a camera. The camera is used to obtain the front-line video in real time and then feed the video back to the master control end. In this process, the lower the delay of the video feedback to the master control end, the better. A deterministic network architecture is constructed so that the master control end can "instantly" receive the video uploaded by the camera.
[0086] This application uses the camera as the video acquisition end, and the port for acquiring the video is called the video receiving end; in most cases, the video receiving end and the video acquisition end are in a one-to-many relationship. Of course, there are also some other architectures, such as a distributed processing architecture. In this case, multiple video receiving ends will be set; in obtaining the transmission performance parameters of the video acquisition end and the transmission performance parameters of the video receiving end, the transmission performance parameters reflect the difficulty of the data transmission process, and the virtual distance between any video acquisition end and the video receiving end is determined according to the transmission performance parameters. The longer the virtual distance, the more difficult the transmission. After processing, all video acquisition ends and video receiving ends are converted into a node-to-node architecture, which is called a video transmission network.
[0087] The video transmission network represents the difficulty of data transmission between various ports. The resource allocation ratio of each video acquisition terminal is determined based on the video transmission network. Then, after obtaining the video, the video acquisition terminal will also identify the video, judge the video itself, and determine the volatility and predictability of the video. The more obvious the volatility and the worse the predictability, the more resources must be invested to ensure the certainty of its data transmission. Accordingly, the resource allocation ratio corresponding to the camera must be increased.
[0088] Finally, the existing transmission resources are allocated according to the adjusted resource allocation ratio, thereby building a deterministic network in which each video acquisition end can transmit data in "real time".
[0089] Among them, the video fluctuation feature and the video prediction feature are both independently extracted by the video acquisition end, the video fluctuation feature is used to characterize the degree of change of each image in the video over time, and the video prediction feature is used to characterize the prediction accuracy of each image in the video; the resource allocation ratio is directly proportional to the video fluctuation feature and inversely proportional to the video prediction feature.
[0090] Embodiment 2:
[0091] Regarding step S100, the steps of obtaining the transmission performance parameters of the video acquisition end and the transmission performance parameters of the video receiving end, and building a video transmission network according to the transmission performance parameters of the video acquisition end, the transmission performance parameters of the video receiving end, and the distance between the video acquisition end and the video receiving end include:
[0092] Querying the transmission performance parameters of the video acquisition end and the transmission performance parameters of the video receiving end in the port filing data; the transmission performance parameters adopt at least one of bandwidth, delay and throughput;
[0093] Input the transmission performance parameters into a preset conversion model to obtain a virtual distance;
[0094] Get the locations of all video acquisition terminals and all video receiving terminals, and create nodes;
[0095] A directed connection line between nodes is determined based on the virtual distance; the direction of the directed connection line is from the video acquisition end to the video receiving end, and the value of the directed connection line is the virtual distance.
[0096] The video acquisition end and the video receiving end will be registered during installation, which is called port registration data. The transmission performance parameters of the video acquisition end and the transmission performance parameters of the video receiving end are queried in the port registration data, and the transmission performance parameters of the two ports that need to interact are input into the preset conversion model to obtain the virtual distance. It should be noted that in the process of calculating the virtual distance, the actual distance can also be introduced as an influencing parameter. However, in the LAN space, the actual distances of each port are of the same magnitude and have little impact. The transmission difficulty is mainly related to the transmission performance parameters of both parties.
[0097] Get the positions of all video capture ends and all video receiving ends, create one-to-one corresponding nodes according to the position relationship, insert directed connection lines between the nodes where the video transmission ends exist, the direction of the directed connection lines is from the video capture end to the video receiving end, and the value of the directed connection lines is the virtual distance.
[0098] Embodiment three:
[0099] Regarding step S200, the step of determining the resource allocation ratio of each video acquisition terminal based on the video transmission network includes:
[0100] Count the values of each directed connection line and calculate the total value;
[0101] Calculate the ratio of the value of each directed link to the total value;
[0102] For any video acquisition end, all the proportions corresponding to the video acquisition end are counted to obtain the final resource allocation ratio.
[0103] The value of each directed connection line is normalized by dividing it by the sum of all values, and the resulting value is in the range of zero to one. Then, for any video acquisition end, all the proportions corresponding to the video acquisition end are counted to represent the overall difficulty of uploading the video, and resources are allocated according to all the proportions, that is, a final resource allocation ratio is determined.
[0104] Embodiment 4:
[0105] In the technical solution of the present invention, the video acquisition end processes the video independently, and the process of extracting the video fluctuation feature and the video prediction feature by the video acquisition end includes the following steps:
[0106] Convert the acquired video into an image set sequence and an audio sequence;
[0107] Identifying the audio sequence, determining an initial separation point, and segmenting the image set sequence based on the initial separation point to obtain an image subset;
[0108] Perform image comparison on each image subset to determine the final separation point;
[0109] Determine the video fluctuation characteristics according to the distribution of the final separation points and the corresponding image comparison results;
[0110] Adjacent separation points to be analyzed are selected from the final separation points, and video prediction features are determined based on images within the adjacent separation points.
[0111] The video includes audio information and image information. The video is separated to obtain an image set sequence and an audio sequence. The audio sequence is identified and the audio is segmented. This segmentation is mainly in the time domain, and audios with similar amplitudes are grouped into one segment. Audio segmentation technology is very common in the prior art and will not be described in detail in this application. After the segmentation is completed, the first moment and the last moment of each audio segment are used as the initial separation points.
[0112] The image set sequence is divided based on the initial separation point to obtain image subsets, and then each image subset is compared for more than half an hour to determine the final separation point; the process of determining the final separation point and the process of determining the initial separation point are actually a gradient process. Audio segmentation is relatively simple, and image comparison is relatively complex. Therefore, this application first recognizes the audio, that is, based on the image comparison, a pre-audio recognition process is introduced.
[0113] Then, the video fluctuation characteristics are determined according to the distribution of the final separation points and the corresponding image comparison results, and then the adjacent separation points to be analyzed are selected from the final separation points, and the video prediction characteristics are determined according to the images in the adjacent separation points.
[0114] Furthermore, the step of identifying the audio sequence, determining the initial separation point, and segmenting the image set sequence based on the initial separation point to obtain the image subset includes:
[0115] Segment the audio sequence according to the audio amplitude to obtain subsequences;
[0116] Input each subsequence into the preset audio recognition model and output the recognized text;
[0117] Merging the subsequences based on the recognized text; wherein, when the recognized text has a dialogue feature, executing the merging process; when the recognized texts belong to the same preset audio media, executing the merging process;
[0118] The first and last moments of each merged subsequence are used as initial separation points;
[0119] The image set sequence is segmented based on the initial separation point to obtain image subsets.
[0120] The audio segmentation process itself is not complicated. The audio sequence is segmented according to the audio amplitude to obtain subsequences. However, in actual applications, the two subsequences may actually be one sequence. For example, in a conversation segment in a video, each person's words will be divided into one segment. In fact, they belong to the same conversation. Based on this, the present application inputs each subsequence into a preset audio recognition model and outputs a recognized text. When the recognized text has conversation features, the two subsequences are merged. When the recognized texts belong to the same preset audio media (belonging to the same BGM), the merging process is performed. Finally, the first and last moments of each merged subsequence are used as initial separation points, and the image set sequence is segmented based on the initial separation points to obtain image subsets.
[0121] It should be noted that the above-mentioned audio segmentation process, audio recognition model and dialogue feature extraction can be completed with the help of existing models, the simplest of which is the AI model. Therefore, the technology for processing audio is mature enough, so this application will not go into details.
[0122] Specifically, the step of performing image comparison on each image subset to determine the final separation point includes:
[0123] For each image subset, the images are read sequentially based on the time sequence and the images are grayed out;
[0124] Read the gray value of each pixel in the image according to a preset pixel reading order, and perform functionization on the image;
[0125] Perform Fourier transform on the functionalized image to obtain a frequency domain image;
[0126] For frequency domain images at adjacent moments, select a region of a preset size in the frequency domain image, compare the selected region, and calculate the image similarity;
[0127] Select images whose image similarity is less than the preset similarity threshold and use the corresponding time point as the final separation point.
[0128] In an example of the technical solution of the present invention, a specific image comparison scheme is provided. The image comparison process is to compare the images in each image subset (the images between two initial separation points). For each image subset, the images are read in sequence based on the time sequence, and the images are grayed so that the color values are simplified to gray values. The gray values are read in sequence from top to bottom and from left to right, the images are functionalized, and the functionalized images are Fourier transformed to obtain a frequency domain image. The closer the position is to the center in the frequency domain image, the lower the corresponding frequency is, and the low frequency corresponds to the pixels with relatively smooth changes in the original image. Based on this feature, for the frequency domain images at adjacent moments, an area of a preset size is selected in the frequency domain image, the selected area is compared, and the image similarity is calculated. When the similarity between the two images is small enough, it means that the difference between them is large, and the corresponding moment is marked as the final separation point (the two parties to the comparison can select one as the final separation point).
[0129] In the above content, the process of selecting an area of a preset size in the frequency domain image is very important. In the frequency domain image, the high-frequency part corresponds to the contour information in the original image. Therefore, the present application can select a high-frequency area in the frequency domain image for comparison. The method of selecting the high-frequency area is to first determine a circle with the origin, and then calculate the complement of the circle in the frequency domain image. The larger the area of the complement, the more content needs to be compared. In the technical solution of the present invention, the area of the area is determined based on the CPU occupancy rate of the video acquisition end. The lower the GPU occupancy rate, the more idle resources the video acquisition end has. At this time, a larger area is needed.
[0130] Embodiment five:
[0131] Video fluctuation characteristics and video prediction characteristics are very important parameters in this application, and the determination process is as follows:
[0132] Regarding the video fluctuation feature, the step of determining the video fluctuation feature according to the distribution of the final separation points and the corresponding image comparison results includes:
[0133] Calculate the distance between each final separation point and the adjacent final separation point and determine the correction factor;
[0134] Read the image similarity corresponding to each final separation point;
[0135] Determine the fluctuation value of each final separation point according to the correction coefficient and the image similarity;
[0136] The fluctuation values of all final separation points are counted as the video fluctuation feature.
[0137] On the premise that the image similarity is known, the distance between two adjacent final separation points is calculated, and the correction coefficient is determined according to the distance. The image similarity at the final separation point is corrected, and the fluctuation value at the final separation point is determined according to the corrected image similarity. Finally, the fluctuation values of all final separation points are counted as the video fluctuation feature.
[0138] Specifically, the calculation process of the video fluctuation feature is:
[0139] Where B is the video fluctuation feature, B i is the fluctuation value of the ith final separation point, N is the total number of final separation points, α and β are the preset correction coefficients, S i is the image similarity at the i-th final separation point, min{d(i, i-1), d(i, i+1)) represents the minimum distance, which is the minimum value of the time difference between the i-th final separation point and the i-1-th final separation point and the time difference between the i+1-th final separation point and the i-th final separation point; when i is 1 and N, there is only one adjacent final separation point, and the time difference is directly read as the minimum value.
[0140] The calculation principle of the video fluctuation feature is not complicated. The image similarity at each final separation point is read. The image at the final separation point is the location where the maximum fluctuation occurs, and the fluctuation value is inversely proportional to the image similarity. On this basis, if the number of final separation points is smaller, it means that the number of locations where the maximum fluctuation occurs is smaller. This application uses the distance between the final separation point and the nearest adjacent separation point to reflect the number of final separation points, and the fluctuation value is inversely proportional to the distance.
[0141] After calculating the fluctuation values of all final separation points, the average of the fluctuation values is calculated as the video fluctuation feature.
[0142] Regarding the video prediction feature, the step of selecting adjacent separation points to be analyzed from the final separation points and determining the video prediction feature according to the images in the adjacent separation points includes:
[0143] Select two adjacent separation points in the final separation points in turn as adjacent separation points to be analyzed;
[0144] Select the image corresponding to the midpoint of adjacent separation points as the reference image;
[0145] Calculating the similarity between each image in adjacent separation points and the reference image, and calculating the image prediction accuracy based on the similarity;
[0146] The image prediction accuracy corresponding to all adjacent separation points is counted and the video prediction features are calculated.
[0147] Adjacent final separation points are read in chronological order as adjacent separation points to be analyzed. There are many images between adjacent separation points. The middle image (the image corresponding to the midpoint of the adjacent separation points) is selected as the reference image. The similarity between the reference image and each image between the adjacent separation points is calculated. The image prediction accuracy can be determined according to the size of the similarity.
[0148] Each pair of adjacent separation points corresponds to an image prediction accuracy. The image prediction accuracy corresponding to all adjacent separation points is counted to obtain the final video prediction accuracy as the video prediction feature.
[0149] Furthermore, the process of determining the image prediction accuracy is as follows:
[0150] Where, T z is the image prediction accuracy, S j is the similarity between the jth image in the adjacent separation point and the reference image, and M is the total number of images in the adjacent separation point;
[0151] The calculation process of the video prediction feature is:
[0152] Where, T v is the video prediction feature, T zn Represents the image prediction accuracy of the nth pair of adjacent separation points.
[0153] The video prediction feature is used to describe predictability. After selecting the reference image, the image similarity between other images and the reference image is calculated using the similarity calculation scheme provided in this application. Then, the mean of the image similarities is calculated and directly used as the image prediction accuracy. Of course, an increasing function can be introduced between the image prediction accuracy and the mean of the similarity to adjust the quantitative relationship.
[0154] Similarly, a video is composed of multiple images. After calculating the image prediction accuracy of each image segment, the mean of the prediction accuracy of all images is calculated to obtain the video prediction feature.
[0155] It should be noted that this application actually mentions a large number of adjacent concepts. A parameter can be considered adjacent to the previous parameter and the next parameter. In practical applications, workers can choose any one of the situations as a standard to make the entire calculation process more regular.
[0156] Specifically, the process of adjusting the resource allocation ratio according to the video fluctuation characteristics and the video prediction characteristics includes:
[0157] Wherein, F′ is the resource allocation ratio after adjustment, c1 and c2 are preset constants, and F is the resource allocation ratio before adjustment.
[0158] The resource allocation ratio is directly proportional to the video fluctuation characteristics and inversely proportional to the video prediction characteristics. It should be noted that the sum of the resource allocation ratios before adjustment is one, but it is not certain after adjustment. The overall resources required may decrease or increase. This also means that in actual situations, resources may occasionally need to be supplemented, and occasionally additional resources are used to complete other tasks.
[0159] Embodiment six:
[0160] Figure 2 The structure diagram of the system for constructing a deterministic network for visual information transmission is shown. In a preferred embodiment of the technical solution of the present invention, a system for constructing a deterministic network for visual information transmission is also provided. The system 10 includes:
[0161] The network construction module 11 is used to obtain the transmission performance parameters of the video acquisition end and the transmission performance parameters of the video receiving end, and to build a video transmission network according to the transmission performance parameters of the video acquisition end, the transmission performance parameters of the video receiving end and the distance between the video acquisition end and the video receiving end;
[0162] An allocation ratio determination module 12, used to determine the resource allocation ratio of each video acquisition terminal based on the video transmission network;
[0163] An allocation ratio adjustment module 13 is used to receive video fluctuation characteristics and video prediction characteristics fed back in real time by a video acquisition end, and adjust the resource allocation ratio according to the video fluctuation characteristics and video prediction characteristics;
[0164] An allocation ratio application module 14, configured to allocate transmission resources based on the adjusted resource allocation ratio;
[0165] Among them, the video fluctuation feature and the video prediction feature are both independently extracted by the video acquisition end, the video fluctuation feature is used to characterize the degree of change of each image in the video over time, and the video prediction feature is used to characterize the prediction accuracy of each image in the video; the resource allocation ratio is directly proportional to the video fluctuation feature and inversely proportional to the video prediction feature.
[0166] Furthermore, the network construction module 11 includes:
[0167] A parameter query unit, used to query the transmission performance parameters of the video acquisition end and the transmission performance parameters of the video receiving end in the port filing data; the transmission performance parameters adopt at least one of bandwidth, delay and throughput;
[0168] A virtual distance generating unit, used for inputting the transmission performance parameter into a preset conversion model to obtain a virtual distance;
[0169] A node creation unit, used to obtain the positions of all video acquisition terminals and all video receiving terminals, and create nodes;
[0170] The node connection unit is used to determine the directed connection line between nodes based on the virtual distance; the direction of the directed connection line is from the video acquisition end to the video receiving end, and the value of the directed connection line is the virtual distance.
[0171] Specifically, the allocation ratio determination module 12 includes:
[0172] A numerical sum calculation unit is used to count the values of each directed connection line and calculate the sum of the values;
[0173] A ratio calculation unit, used to calculate the ratio of the value of each directed link to the total value;
[0174] The ratio statistics unit is used to count all the ratios corresponding to any video acquisition end to obtain the final resource allocation ratio.
[0175] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A method for constructing a deterministic network for visual information transmission, characterized in that: The method comprises: Acquire the transmission performance parameters of the video acquisition end and the transmission performance parameters of the video receiving end, and build a video transmission network according to the transmission performance parameters of the video acquisition end, the transmission performance parameters of the video receiving end and the distance between the video acquisition end and the video receiving end; Determine the resource allocation ratio of each video acquisition terminal based on the video transmission network; Receive video fluctuation characteristics and video prediction characteristics fed back in real time by the video acquisition end, and adjust the resource allocation ratio according to the video fluctuation characteristics and video prediction characteristics; Allocating transmission resources based on the adjusted resource allocation ratio; Among them, the video fluctuation feature and the video prediction feature are both independently extracted by the video acquisition end, the video fluctuation feature is used to characterize the degree of change of each image in the video over time, and the video prediction feature is used to characterize the prediction accuracy of each image in the video; the resource allocation ratio is directly proportional to the video fluctuation feature and inversely proportional to the video prediction feature.
2. The method for constructing a deterministic network for visual information transmission according to claim 1, characterized in that: The steps of obtaining the transmission performance parameters of the video acquisition end and the transmission performance parameters of the video receiving end, and building a video transmission network according to the transmission performance parameters of the video acquisition end, the transmission performance parameters of the video receiving end and the distance between the video acquisition end and the video receiving end include: Querying the transmission performance parameters of the video acquisition end and the transmission performance parameters of the video receiving end in the port filing data; the transmission performance parameters adopt at least one of bandwidth, delay and throughput; Input the transmission performance parameters into a preset conversion model to obtain a virtual distance; Get the locations of all video acquisition terminals and all video receiving terminals, and create nodes; A directed connection line between nodes is determined based on the virtual distance; the direction of the directed connection line is from the video acquisition end to the video receiving end, and the value of the directed connection line is the virtual distance.
3. The method for constructing a deterministic network for visual information transmission according to claim 2, characterized in that: The step of determining the resource allocation ratio of each video acquisition terminal based on the video transmission network includes: Count the values of each directed connection line and calculate the total value; Calculate the ratio of the value of each directed link to the total value; For any video acquisition end, all the proportions corresponding to the video acquisition end are counted to obtain the final resource allocation ratio.
4. The method for constructing a deterministic network for visual information transmission according to claim 1, characterized in that: The process of extracting video fluctuation features and video prediction features at the video acquisition end includes: Convert the acquired video into an image set sequence and an audio sequence; Identifying the audio sequence, determining an initial separation point, and segmenting the image set sequence based on the initial separation point to obtain an image subset; Perform image comparison on each image subset to determine the final separation point; Determine the video fluctuation characteristics according to the distribution of the final separation points and the corresponding image comparison results; Adjacent separation points to be analyzed are selected from the final separation points, and video prediction features are determined based on images within the adjacent separation points.
5. The method for constructing a deterministic network for visual information transmission according to claim 4, characterized in that: The steps of identifying the audio sequence, determining the initial separation point, and segmenting the image set sequence based on the initial separation point to obtain the image subset include: Segment the audio sequence according to the audio amplitude to obtain subsequences; Input each subsequence into the preset audio recognition model and output the recognized text; Merging the subsequences based on the recognized text; wherein, when the recognized text has a dialogue feature, executing the merging process; when the recognized texts belong to the same preset audio media, executing the merging process; The first and last moments of each merged subsequence are used as initial separation points; The image set sequence is segmented based on the initial separation point to obtain image subsets.
6. The method for constructing a deterministic network for visual information transmission according to claim 5, characterized in that: The step of performing image comparison on each image subset to determine the final separation point comprises: For each image subset, the images are read sequentially based on the time sequence and grayscaled; Read the gray value of each pixel in the image according to a preset pixel reading order, and perform functionization on the image; Perform Fourier transform on the functionalized image to obtain a frequency domain image; For frequency domain images at adjacent moments, a region of a preset size is selected in the frequency domain image, the selected region is compared, and the image similarity is calculated; Select images whose image similarity is less than a preset similarity threshold, and use the corresponding time point as the final separation point; In the process of selecting a region of a preset size in the frequency domain image, the area of the region is inversely proportional to the CPU occupancy rate of the video acquisition end.
7. The method for constructing a deterministic network for visual information transmission according to claim 6, characterized in that: The step of determining the video fluctuation characteristics according to the distribution of the final separation points and the corresponding image comparison results comprises: Calculate the distance between each final separation point and the adjacent final separation point and determine the correction factor; Read the image similarity corresponding to each final separation point; Determine the fluctuation value of each final separation point according to the correction coefficient and the image similarity; The fluctuation values of all final separation points are counted as the video fluctuation feature.
8. The method for constructing a deterministic network for visual information transmission according to claim 7, characterized in that: The calculation process of the video fluctuation feature is as follows: Where B is the video fluctuation feature, B i is the fluctuation value of the ith final separation point, N is the total number of final separation points, α and β are the preset correction coefficients, S i is the image similarity at the i-th final separation point, min{d(i,i-1),d(i,i+1)} represents the minimum distance, which is the minimum value of the time difference between the i-th final separation point and the i-1-th final separation point and the time difference between the i+1-th final separation point and the i-th final separation point; when i is 1 and N, there is only one adjacent final separation point, and the time difference is directly read as the minimum value.
9. The method for constructing a deterministic network for visual information transmission according to claim 8, characterized in that: The step of selecting adjacent division points to be analyzed from the final division points and determining the video prediction features according to the images in the adjacent division points comprises: Select two adjacent separation points in the final separation points in turn as adjacent separation points to be analyzed; Select the image corresponding to the midpoint of adjacent separation points as the reference image; Calculating the similarity between each image in adjacent separation points and the reference image, and calculating the image prediction accuracy based on the similarity; The image prediction accuracy corresponding to all adjacent separation points is counted and the video prediction features are calculated.
10. The method for constructing a deterministic network for visual information transmission according to claim 9, characterized in that: The process of determining the image prediction accuracy is as follows: Where, T z is the image prediction accuracy, S j is the similarity between the jth image in the adjacent separation point and the reference image, and M is the total number of images in the adjacent separation point; The calculation process of the video prediction feature is: Where, T v is the video prediction feature, T zn Represents the image prediction accuracy of the nth pair of adjacent separation points.
11. The method for constructing a deterministic network for visual information transmission according to claim 10, characterized in that: The process of adjusting the resource allocation ratio according to the video fluctuation characteristics and the video prediction characteristics includes: In the formula, F ′ is the resource allocation ratio after adjustment, c1 and c2 are preset constants, and F is the resource allocation ratio before adjustment.
12. A system for constructing a deterministic network for visual information transmission, characterized in that: The system comprises: A network construction module is used to obtain the transmission performance parameters of the video acquisition end and the transmission performance parameters of the video receiving end, and to build a video transmission network according to the transmission performance parameters of the video acquisition end, the transmission performance parameters of the video receiving end and the distance between the video acquisition end and the video receiving end; An allocation ratio determination module, used to determine the resource allocation ratio of each video acquisition terminal based on the video transmission network; An allocation ratio adjustment module is used to receive video fluctuation characteristics and video prediction characteristics fed back in real time by a video acquisition end, and adjust the resource allocation ratio according to the video fluctuation characteristics and video prediction characteristics; An allocation ratio application module, used for allocating transmission resources based on the adjusted resource allocation ratio; Among them, the video fluctuation feature and the video prediction feature are both independently extracted by the video acquisition end, the video fluctuation feature is used to characterize the degree of change of each image in the video over time, and the video prediction feature is used to characterize the prediction accuracy of each image in the video; the resource allocation ratio is directly proportional to the video fluctuation feature and inversely proportional to the video prediction feature.
13. The system for constructing a visual information transmission deterministic network according to claim 12, characterized in that: The network building module includes: A parameter query unit, used to query the transmission performance parameters of the video acquisition end and the transmission performance parameters of the video receiving end in the port filing data; the transmission performance parameters adopt at least one of bandwidth, delay and throughput; A virtual distance generating unit, used for inputting the transmission performance parameter into a preset conversion model to obtain a virtual distance; A node creation unit, used to obtain the positions of all video acquisition terminals and all video receiving terminals, and create nodes; The node connection unit is used to determine the directed connection line between nodes based on the virtual distance; the direction of the directed connection line is from the video acquisition end to the video receiving end, and the value of the directed connection line is the virtual distance.
14. The system for constructing a deterministic network for visual information transmission according to claim 12, characterized in that: The allocation ratio determination module includes: A numerical sum calculation unit is used to count the values of each directed connection line and calculate the sum of the values; A ratio calculation unit, used to calculate the ratio of the value of each directed link to the total value; The ratio statistics unit is used to count all the ratios corresponding to any video acquisition end to obtain the final resource allocation ratio.