Image identification and content optimization method for short video platform projection stream

By performing frame-by-frame processing and user feature analysis on short videos, the best cover image was selected and the area was optimized. This solved the problem of cover image design relying on human experience in traditional streaming methods, and improved the visual appeal and user interaction of short videos.

CN121708522APending Publication Date: 2026-03-20FEIKE WANGHONG (HANGZHOU) TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511785564.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-01
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Traditional short video distribution methods lack a deep understanding of video content, and cover image design relies on human experience, making it impossible to achieve personalized optimization, resulting in low content reach and low user conversion efficiency.

Method used

By segmenting short videos into frames and extracting image feature data, combined with multi-dimensional user feature analysis, the best cover images that fit different user groups are selected, and saliency analysis and regional optimization design are carried out to reasonably arrange the cover text.

Benefits of technology

It enables personalized cover image selection for different user groups, enhances the visual appeal and conversion rate of short videos, increases click-through rate and user interaction, and optimizes the accuracy and effectiveness of the advertising strategy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121708522A_ABST
    Figure CN121708522A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer vision, and discloses an image recognition and content optimization method for short video platform streaming. The method comprises the following steps: framing a short video, extracting each frame of video image, and extracting image feature data; performing feature division on the traffic casting target users to obtain common feature data of different user groups; performing feature matching on the image feature data and the common feature data, and screening out an optimal cover image matched with different user groups; performing significance analysis on each optimal cover image, positioning a visual attention area, and arranging the designed cover copywriting in the visual attention area; performing image quality optimization on each optimal cover image in sequence, and performing differentiated short video delivery for different user groups; according to the invention, an accurate delivery strategy based on content attributes and user features can be realized, and the touch effect and the propagation value of short video content are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and more specifically, to a method for image recognition and content optimization for short video platform streaming. Background Technology

[0002] With the widespread adoption of smartphones and mobile internet, short videos have become one of the main ways for users to obtain information and entertainment. Along with the rapid expansion of the user base, the number of short video creators has also increased significantly, resulting in an explosive growth in short video content and the formation of a highly competitive content ecosystem. Against this backdrop, traffic delivery has become a crucial element for platforms and content creators to accurately reach users, increase exposure, and enhance user interaction. Through precise traffic delivery strategies, suitable content can be recommended to potentially interested users, achieving efficient content dissemination.

[0003] However, with the diversification of user needs and the massive supply of content, how to achieve accurate and efficient traffic delivery has become a major challenge for platforms and creators. Traditional traffic delivery methods have many shortcomings: on the one hand, they mainly rely on users' historical behavior data and simple tag matching, lacking a deep understanding of the video content itself; on the other hand, the cover image, as the first visual entry point for users to decide whether to click to watch, often relies on human experience in its design, lacking data support and consideration of user group differences, making it difficult to achieve personalized cover optimization for different user groups, failing to meet the needs of refined operation, resulting in low content reach and low user conversion efficiency, which in turn restricts the maximization of the platform's content value and the continuous improvement of user experience.

[0004] In view of this, the present invention proposes an image recognition and content optimization method for short video platform streaming to solve the above problems. Summary of the Invention

[0005] To overcome the aforementioned deficiencies of the prior art and achieve the above objectives, the present invention provides the following technical solution: a method for image recognition and content optimization for short video platform streaming, comprising: Step S1: Perform frame segmentation on the short video and extract the video image sequence; Step S2: Perform subject detection on each frame of the video image sequence in sequence, identify the corresponding subject content, and extract image feature data based on the subject content; Step S3: Perform multi-dimensional feature segmentation on the target users to obtain common feature data corresponding to different user groups; Step S4: Perform feature matching between image feature data and common feature data, and select the best cover image that matches different user groups from the video images based on the matching results; Step S5: Perform a salience analysis on each best cover image in turn, locate the visual attention area, and place the designed cover text in the visual attention area of ​​each best cover image; Step S6: Optimize the image quality of each best cover image in turn, and based on the optimized best cover image, implement differentiated short video delivery for different user groups.

[0006] Furthermore, methods for identifying the main content corresponding to a video image include: The video image is input into the trained subject detection model, and the subject detection result is output, which includes the subject localization rectangle and the subject type label. The video image is cropped based on the main subject positioning rectangle to obtain the main subject region within the video image; the main subject region is combined with the main subject type tag to form the main subject content. Methods for extracting image feature data based on main content include: Color analysis is performed on the main content to extract color distribution and main color tone; composition and style analysis is performed on the video images to extract visual style tags, subject position, and image balance; the subject type tags, color distribution, main color tone, visual style tags, subject position, and image balance are integrated to form image feature data.

[0007] Furthermore, methods for extracting color distribution include: Convert the main area from RGB color space to HSV color space, obtain all pixels in the main area, and represent each pixel with the corresponding HSV color value; K-means clustering is used to cluster all pixels to obtain multiple color clusters. The number of pixels in each color cluster is counted and labeled as the pixel count. The sum of all pixel counts is calculated to obtain the total number of pixels. For each color cluster, the mean of the HSV color values ​​of all pixels is calculated to obtain the color center. The ratio of the number of pixels in each cluster to the total number of pixels is used as the color proportion of the corresponding color center. Based on all color centers and their corresponding color proportions, the color distribution of the main region is constructed. Methods for extracting image balance include: The video image is divided into nine equal regions according to the rule of thirds. The corresponding color tone of each region is extracted sequentially, and the brightness value of the corresponding color tone in each region is obtained. Each brightness value is normalized to obtain a standard brightness value. Based on each standard brightness value, the color of each region is calculated. The area of ​​the overlapping portion between each region and the main body region is calculated and marked as the intersection area. The image width and height of the video image are obtained, and the area of ​​each region is calculated. The ratio of each intersection area to the corresponding region area is calculated to obtain the area proportion. The product of the region color, area proportion, and a preset weight for each region is calculated to obtain the visual weight of each region. The coordinates of the center of each equally divided region are calculated based on the image width and image height. The coordinates of the center of all regions are weighted and averaged to obtain the coordinates of the visual centroid. The coordinates of the geometric center of the video image are calculated based on the image width and image height, and the Euclidean distance is used to measure the actual deviation between the coordinates of the visual centroid and the geometric center. The maximum deviation is calculated, and the image balance is calculated based on the actual deviation and the maximum deviation.

[0008] Furthermore, methods for obtaining common characteristic data corresponding to different user groups include: Acquire the historical short videos corresponding to each target user, as well as the user behavior data corresponding to each historical short video; among which, the user behavior data includes viewing time and interaction methods, including likes, comments and shares; Obtain the total duration of each historical short video, calculate the ratio of each viewing duration to the corresponding total video duration, and obtain the video completion rate of each historical short video; compare each video completion rate with a preset completion threshold, and delete historical short videos with a completion rate less than the completion threshold. Based on video completion rate and interaction method, calculate the interest weight corresponding to each historical short video; obtain the cover image of each historical short video and mark it as a historical cover image; extract the corresponding image feature data from each historical cover image and mark it as historical feature data; based on the interest weight, calculate the weighted average of the historical feature data corresponding to the historical short videos of the same target user to obtain the interest feature data corresponding to each target user. Based on the interest feature data, the K-means clustering method is used to cluster all target users of the traffic delivery, resulting in multiple user clusters, which are then labeled as user groups. For each user group, the mean of the interest feature data corresponding to all target users of the traffic delivery is calculated to obtain the common feature data corresponding to each user group.

[0009] Furthermore, methods for calculating the interest weights corresponding to historical short videos include: Sentiment analysis is performed on comments in interactive content to obtain the emotional state of the target users. A weight set is preset, which includes weight coefficients for likes, shares, and different emotional states. The viewing time of historical short videos is obtained, and the difference between the current time and the viewing time is calculated to obtain the time difference. A time decay function is preset, and the time difference is substituted into the time decay function to obtain the time weight. The weight coefficients corresponding to likes, shares, and emotional states are added in sequence to obtain the behavior weight. The behavior weight, video completion rate, and time weight are multiplied in sequence to obtain the interest weight.

[0010] Furthermore, methods for feature matching between image feature data and common feature data include: Based on all image feature data and common feature data, construct multiple sets of different feature data; for each set of common feature data, calculate the degree of difference between each data point in the image feature data and the corresponding data point in the common feature data. For each type of data in the image feature data, a corresponding data weight is dynamically assigned; based on the data weight, all the degree of difference corresponding to the same feature data set are weighted and summed to obtain the overall matching degree corresponding to each set of feature data.

[0011] Furthermore, methods for dynamically assigning data weights include: For each user group, the local dispersion of each type of data in the image feature data is calculated sequentially; the mean of the local dispersion of the same data type in the corresponding image feature data is calculated to obtain the average dispersion of each type of data in the image feature data, and the reciprocal of the average dispersion is used as the corresponding data weight.

[0012] Furthermore, methods for selecting the best cover images from video images to match different user groups include: The overall matching degree corresponding to the same common feature data is compared, and the video image corresponding to the overall matching degree with the highest value is used as the candidate cover image for the corresponding user group; the overall matching degree corresponding to each candidate cover image is compared with the preset matching threshold. If the overall matching degree is greater than or equal to the matching threshold, the corresponding candidate cover image will be used as the best cover image for the corresponding user group. If the overall matching degree is less than the matching threshold, the corresponding user group will be marked as the group to be analyzed; Based on the common feature data of the group to be analyzed, a style transfer model is used to sequentially perform style transfer on all video images to obtain transferred images. The corresponding image feature data is extracted from each transferred image and marked as transferred feature data. Each set of transferred feature data is matched with the common feature data of the group to be analyzed to obtain the overall matching degree corresponding to each transferred image and marked as the transfer matching degree. All transfer matching degrees are compared, and the transferred image corresponding to the transfer matching degree with the largest value is used as the best cover image for the corresponding group to be analyzed.

[0013] Furthermore, methods for locating visually attenuated areas include: A pre-trained saliency analysis model is used to perform saliency analysis on each best cover image in turn, obtaining the set of salient points corresponding to each best cover image. The set of salient points includes multiple salient parameters corresponding to the salient points, including position coordinates and saliency intensity. According to the cover text, the region parameters of the corresponding text area are determined. Based on the set of salient points and region parameters corresponding to each best cover image, the text area corresponding to each best cover image is constructed in turn. In each best cover image, the corresponding text area is mapped sequentially. The area of ​​the overlapping part between each text area and the main body area is calculated and marked as the overlapping area. Each overlapping area is compared with a preset area threshold, and salient points with overlapping areas greater than the area threshold are deleted from all salient point sets. The deleted salient point set is marked as the candidate point set. Based on the candidate point set corresponding to each best cover image and the area parameters, the text area corresponding to each best cover image is reconstructed and marked as the reconstructed area. The corresponding reconstructed area is added sequentially to each best cover image, and the cover text is added to each reconstructed area to obtain the text cover image. The image balance of each cover image is extracted sequentially. Based on a preset set of proportions, the image balance of each cover image and the salience intensity of the corresponding salient points are weighted and summed to obtain the layout quality of each cover image. All layout quality of the same cover image are compared, and the reconstructed area with the highest layout quality is taken as the visual attention area of ​​the corresponding cover image.

[0014] Furthermore, the method for sequentially optimizing the image quality of each best cover image includes: The connected component labeling method is used to perform connectivity analysis on the salient points corresponding to each best cover image, thereby obtaining the salient regions corresponding to each best cover image; In each best cover image, the region consisting of pixels that are not salient points is marked as the background region; an adaptive histogram equalization method is used to dynamically adjust the contrast of each salient region in each best cover image; and an edge-preserving filtering algorithm is applied to perform global consistency processing on each best cover image to complete the image quality optimization of each best cover image.

[0015] The technical effects and advantages of the image recognition and content optimization method for short video platform streaming according to the present invention are as follows: By performing frame-by-frame processing, subject detection, and image feature extraction on short videos, combined with multi-dimensional feature analysis of target users, personalized optimal cover image selection can be achieved for different user groups, thereby enhancing the targeting and attractiveness of short video cover images. After determining the optimal cover image, saliency analysis and regional optimization design methods can automatically locate visual attention areas and reasonably arrange cover text, fully satisfying the information preferences of different user groups and effectively improving the overall visual appeal of short videos. Image quality optimization for each optimal cover image involves local detail enhancement and global consistency processing, which not only improves the visual refinement of the cover image but also ensures clear readability in complex backgrounds, thereby improving the conversion rate of short video delivery. A precise delivery strategy based on content attributes and user characteristics can be implemented. Compared with traditional methods that rely on behavioral data and tag matching, this approach can more accurately grasp user needs, thereby significantly improving the click-through rate and user interaction of video content, optimizing the accuracy and effectiveness of short video delivery, and ultimately enhancing the reach and dissemination value of short video content. Attached Figure Description

[0016] Figure 1 This is a flowchart of an image recognition and content optimization method for short video platform streaming according to Embodiment 1 of the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Example 1:

[0018] Please see Figure 1 As shown in this embodiment, an image recognition and content optimization method for short video platform streaming includes: Step S1: Perform frame segmentation on the short video and extract the video image sequence.

[0019] The short video is decoded using a video decoder (such as a decoding module based on FFmpeg or OpenCV). Each frame of the short video is read sequentially and converted into a static image format to form a video image sequence. The video image sequence contains the video images corresponding to all frames in the short video.

[0020] Step S2: Perform subject detection on each frame of the video image sequence in sequence, identify the corresponding subject content, and extract image feature data based on the subject content.

[0021] Methods for identifying the main content corresponding to a video image include: The video image is input into a trained subject detection model (such as YOLOv5, Faster R-CNN, etc.), and the subject detection result is output, which includes the subject localization rectangle and the subject type label.

[0022] The video image is cropped based on the main positioning rectangle to obtain the main area within the video image; the main area is combined with the main type tag to form the main content.

[0023] The subject positioning rectangle is used to mark the area where the subject is located in the video image, and it is usually represented by four coordinate values ​​(i.e., the coordinates of the top left vertex and the bottom right vertex). The subject type label is a numeric label that represents the subject type. Different subject types correspond to different numeric labels. Subject types include people, products, pets, etc.

[0024] Methods for extracting image feature data based on main content include: Color analysis is performed on the main content to extract color distribution and main color tone; composition and style analysis is performed on the video images to extract visual style tags, subject position, and image balance; the subject type tags, color distribution, main color tone, visual style tags, subject position, and image balance are integrated to form image feature data.

[0025] Methods for extracting color distribution include: The main area is converted from RGB color space to HSV color space to improve the ability to express color features and robustness; all pixels in the main area are obtained and each pixel is represented by its corresponding HSV color value; K-means clustering is used to cluster all pixels to obtain multiple color clusters. The number of pixels in each color cluster is counted and labeled as the pixel count. The sum of all pixel counts is calculated to obtain the total number of pixels. For each color cluster, the mean of the HSV color values ​​of all pixels is calculated to obtain the color center. The ratio of the number of pixels in each cluster to the total number of pixels is used as the color proportion of the corresponding color center. Based on all color centers and their corresponding color proportions, the color distribution of the main region is constructed. It should be noted that the methods for converting RGB color space to HSV color space and the K-means clustering method are existing technologies, and the specific processes will not be elaborated on here.

[0026] Methods for extracting the main color tone include: The color proportion of each color center is used as a weighting coefficient, and the weighted sum of all color centers is obtained to obtain the main color tone.

[0027] Methods for extracting visual style tags include: Video images are input into a trained style classification model, which outputs visual style labels. These labels are numerical representations of visual styles, with each style having a unique label. Examples of visual styles include retro, fresh, and technological. The style classification model is a convolutional neural network (CNN), an extension of deep neural networks, primarily consisting of an input layer, multiple convolutional layers, pooling layers (downsampling layers), fully connected layers, and an output layer. Each convolutional layer comprises multiple convolutional kernels (filters), each sliding with the input data to extract local features. Weight parameters are included in the convolution operation to learn the importance of features. Pooling layers typically follow convolutional layers to reduce the size of feature maps, preserving key information while reducing computational complexity. Fully connected layers follow the convolutional and pooling layers to map high-dimensional features to the output space. In the convolutional and fully connected layers, each neuron typically applies an activation function, introducing non-linearity to enhance the model's ability to express complex image patterns and its generalization capabilities. The specific training process of the style classification model includes: Multiple video images are pre-collected, and each video image is labeled with a visual style tag. The labeled video images are divided into a training set and a test set, with 70% of the video images used as the training set and 30% used as the test set. The style classification model is trained using the training set and tested using the test set. A preset error threshold is set, and the style classification model is output when the mean prediction error of all video images in the test set is less than the error threshold. The mean prediction error is calculated using the average cross-entropy loss, and the error threshold is preset according to the accuracy required by the style classification model.

[0028] Methods for extracting the main body location include: The video image is divided into nine equal regions according to the rule of thirds, that is, the video image is divided into three equal parts both horizontally and vertically. The nine equal regions are: upper left region, middle left region, lower left region, upper middle region, center region, lower middle region, upper right region, middle right region, and lower right region. Based on the subject positioning rectangle, the center coordinates of the subject are calculated (i.e., the center coordinates of the subject positioning rectangle). Based on the center coordinates of the subject, the equal regions to which the subject belongs are determined and used as the subject position.

[0029] Methods for extracting image balance include: For each equally divided region, the corresponding region tone is extracted sequentially, and combined with the main body region, the visual weight corresponding to each equally divided region is calculated; the method for extracting the region tone is the same as the method for extracting the main body tone; the image width and image height corresponding to the video image are obtained, and the region center coordinates of each equally divided region are calculated based on the image width and image height; Based on visual weight, a weighted average of the center coordinates of all regions is calculated to obtain the visual centroid coordinates; based on the image width and image height, the geometric center coordinates of the video image are calculated (the horizontal coordinate of the geometric center coordinates is half the image width, and the vertical coordinate is half the image height), and Euclidean distance is used to measure the actual deviation between the visual centroid coordinates and the geometric center coordinates; the maximum deviation is calculated, and the image balance is calculated based on the actual deviation and the maximum deviation.

[0030] The method for calculating visual weight is as follows: Obtain the brightness value in the hue of the corresponding region for each equally divided region; normalize each brightness value to obtain a standard brightness value; calculate the difference between the standard brightness value and the standard brightness value to obtain the region color of each equally divided region; calculate the area of ​​the overlapping part between each equally divided region and the main body region, and mark it as the intersection area; calculate the region area of ​​each equally divided region based on the image width and image height; calculate the ratio of each intersection area to the corresponding region area to obtain the area proportion; calculate the product of the region color, area proportion, and preset weight weight to obtain the visual weight. It should be noted that the method for calculating the intersection area is existing technology, and the specific process will not be elaborated upon here; the weight weight is preset by those skilled in the art based on the actual situation.

[0031] The expression for the coordinates of the region center is: ; In the formula, The x-coordinate of the region's center coordinates. The ordinate of the region's center coordinates. This is the row index for the corresponding equally divided region. , For the column index of the corresponding equally divided region, , Image width, The height is the image height; where the row index represents the vertical position of the divided region, and the column index represents the horizontal position of the divided region; for example, , Then the corresponding equally divided region is the upper-middle region. , Then the corresponding equally divided area is the center area.

[0032] The maximum deviation is the diagonal length of the video image. The method to calculate the maximum deviation is to add the square of the image width to the square of the image height, and then take the square root to obtain the maximum deviation.

[0033] The method for calculating image balance is as follows: calculate the ratio of the deviation distance to the maximum deviation distance to obtain the degree of deviation; take the difference between the deviation distance and the degree of deviation as the image balance.

[0034] Step S3: Perform multi-dimensional feature segmentation on the target users to obtain common feature data corresponding to different user groups.

[0035] Methods for obtaining common feature data corresponding to different user groups include: By using the user behavior log system built into the short video platform, we can obtain the historical short videos corresponding to each target user and the user behavior data corresponding to each historical short video. Among them, the target users for content delivery refer to the users that the short video platform expects to reach when delivering content to the current short video. They are usually generated by the recommendation algorithm built into the short video platform. The current short video is the short video that needs to be optimized. Historical short videos are short videos of the same type as the current short video that the target users watched at historical moments, such as those in the same categories as beauty, games, and fitness. User behavior data includes viewing time and interaction methods. Viewing time is the actual time that the target users stayed when watching historical short videos. Interaction methods include liking, commenting, and sharing. The total duration of each historical short video is obtained, and the ratio of each viewing duration to the corresponding total video duration is calculated to obtain the video completion rate of each historical short video. The completion rate of each video is compared with a preset completion threshold, and historical short videos with a completion rate less than the completion threshold are deleted. The completion threshold is preset by those skilled in the art based on the actual situation. Based on video completion rate and interaction methods, the interest weight corresponding to each historical short video is calculated to measure the overall interest of the target users in different historical short videos; the cover image of each historical short video is obtained and marked as a historical cover image; the corresponding image feature data is extracted from each historical cover image and marked as historical feature data; based on the interest weight, the historical feature data corresponding to the historical short videos of the same target user are weighted and averaged to obtain the interest feature data corresponding to each target user; Based on the interest feature data, the K-means clustering method is used to cluster all target users of the traffic delivery, resulting in multiple user clusters, which are then labeled as user groups. For each user group, the mean of the interest feature data corresponding to all target users of the traffic delivery is calculated to obtain the common feature data corresponding to each user group.

[0036] Methods for calculating the interest weights corresponding to historical short videos include: Sentiment analysis is performed on comments in interactive activities to obtain the emotional state of the target users. Emotional states include positive, negative, and neutral emotions. The methods for obtaining emotional states are existing technologies, such as dictionary-based methods and machine learning methods; the specific process will not be elaborated upon here. A preset weight set is used, including weight coefficients for likes, shares, and different emotional states, which are pre-set by those skilled in the art based on actual conditions. Specifically, the weight coefficient for positive emotions is greater than that for neutral emotions, and the weight coefficient for neutral emotions is greater than that for negative emotions. Obtain the viewing time corresponding to the historical short videos (i.e., the specific time when the target user started watching the historical short videos), and calculate the difference between the current time and the viewing time to obtain the time difference; preset a time decay function, substitute the time difference into the time decay function to obtain the time weight; add the weight coefficients corresponding to likes, shares and emotional states in sequence to obtain the behavior weight; multiply the behavior weight, video completion rate and time weight in sequence to obtain the interest weight; The time decay function is a function that monotonically decreases as the time difference increases. It is used to quantify the impact of the distance between the viewing time and the current time on the time weight. The specific form can be an exponential decay function, a linear decay function, etc., which can be pre-designed by those skilled in the art according to the actual situation.

[0037] Step S4: Perform feature matching between image feature data and common feature data, and select the best cover image from video images that suits different user groups based on the matching results.

[0038] Methods for feature matching between image feature data and common feature data include: Based on all image feature data and common feature data, multiple different feature data sets are constructed. Each feature data set contains a set of image feature data and a set of common feature data. For each set of common feature data, the degree of difference between each data point in the image feature data and the corresponding data point in the common feature data is calculated in turn. For each type of data in the image feature data, a corresponding data weight is dynamically assigned; based on the data weight, all the degree of difference corresponding to the same feature data set are weighted and summed to obtain the overall matching degree corresponding to each set of feature data.

[0039] It should be noted that when calculating the degree of difference, different numerical labels are set for different subject positions and marked as position labels. Since the subject type label, visual style label, and position label are all discrete data, they need to be converted into vector form first (such as one-hot encoding, embedded vector, etc.) and then the degree of difference is calculated using cosine similarity (i.e., the degree of difference is the difference between one and cosine similarity). The degree of difference corresponding to color distribution is calculated using EMD or weighted histogram distance. The degree of difference corresponding to subject color tone is also calculated using cosine similarity. As the image balance is continuous data, the degree of difference is directly measured by the absolute difference.

[0040] Methods for dynamically assigning data weights include: For each user group, the local dispersion of each type of data in the image feature data is calculated sequentially; the mean of the local dispersion of the same data type in the corresponding image feature data is calculated to obtain the average dispersion of each type of data in the image feature data, and the reciprocal of the average dispersion is used as the corresponding data weight, thereby realizing the dynamic weighted allocation of the image feature data. The method for calculating local dispersion is as follows: Each type of data within the image feature data is labeled as a sample. For all interest feature data within a user group, samples of the same type are grouped into a sample set. The difference between every two samples in each sample set is calculated and labeled as the sample difference. For each sample set, the sample differences corresponding to the same samples are grouped into a difference set. The sample differences in each difference set are sorted from smallest to largest, and the top-ranked differences are selected. The difference between the samples is used as the adjacency distance of the corresponding samples. The value is an integer greater than 1; the adjacency distances of the same sample are averaged to obtain the average adjacency distance; the average adjacency distances of the same sample set are averaged to obtain the local dispersion of each sample set; It should be noted that, to prevent errors in calculating the reciprocal of the average dispersion due to a zero average dispersion, a very small positive number should be added to the average dispersion first. This embodiment preferably uses this. This ensures that the denominator is non-zero.

[0041] Methods for selecting the best cover images from video images to match different user groups include: The overall matching degree corresponding to the same common feature data is compared, and the video image corresponding to the overall matching degree with the largest value is used as the candidate cover image for the corresponding user group; the overall matching degree corresponding to each candidate cover image is compared with a preset matching threshold, which is preset by those skilled in the art according to the actual situation; If the overall matching degree is greater than or equal to the matching threshold, the corresponding candidate cover image will be used as the best cover image for the corresponding user group. If the overall matching degree is less than the matching threshold, the corresponding user group will be marked as the group to be analyzed; Based on the common feature data of the group to be analyzed, style transfer models (such as AdaIN, Fast StyleTransfer, etc.) are used to perform style transfer on all video images in sequence to obtain transferred images; the corresponding image feature data are extracted from each transferred image in sequence and marked as transfer feature data; each set of transfer feature data is matched with the common feature data of the group to be analyzed to obtain the overall matching degree corresponding to each transferred image and marked as the transfer matching degree; all transfer matching degrees are compared and the transferred image corresponding to the transfer matching degree with the largest value is used as the best cover image for the corresponding group to be analyzed.

[0042] Step S5: Perform a salience analysis on each best cover image in turn, locate the visual attention area, and place the designed cover text within the visual attention area of ​​each best cover image.

[0043] Methods for locating visual attention areas include: Pre-trained saliency analysis models (such as U²-Net, PoolNet, DeepGaze II, etc.) are used to perform saliency analysis on each best cover image sequentially, obtaining a set of salient points corresponding to each best cover image. The set of salient points includes multiple salient parameters corresponding to each salient point, including location coordinates and saliency intensity. Salient points are visual hotspots, representing the specific locations where visual attention is most easily attracted. Saliency intensity reflects the visual saliency of the salient point; the higher the value, the more easily the salient point attracts the user's attention. The value range is... ; Based on the cover text designed by the short video creator, determine the area parameters of the corresponding text area, including the area height and area width. The area parameters are determined by the short video creator based on the designed cover text and practical experience. Based on the set of salient points corresponding to each best cover image and the area parameters, construct the text area corresponding to each best cover image in turn, with each text area corresponding to a salient point in the set of salient points. In each best cover image, the corresponding text area is mapped sequentially, the area of ​​the overlapping part between each text area and the main body area is calculated and marked as the overlapping area; each overlapping area is compared with a preset area threshold, and salient points in all salient point sets whose overlapping area is greater than the area threshold are deleted. The area threshold is set by those skilled in the art according to the actual situation. The set of salient points after deletion is marked as the candidate point set. Based on the candidate point set and region parameters corresponding to each best cover image, the text region corresponding to each best cover image is reconstructed and marked as the reconstructed region. The corresponding reconstructed region is added to each best cover image in sequence, and the cover text is added to each reconstructed region to obtain the text cover image. Each text cover image includes one reconstructed region. The image balance of each cover image is extracted sequentially. Based on a preset set of proportions, the image balance of each cover image and the salience intensity of the corresponding salient points are weighted and summed to obtain the layout quality of each cover image. All layout quality of the same cover image are compared, and the reconstructed area with the highest layout quality is taken as the visual attention area of ​​the corresponding cover image. The set of proportions includes the proportion coefficients corresponding to the image balance and salience intensity, which are preset by those skilled in the art according to the actual situation.

[0044] Step S6: Optimize the image quality of each best cover image in turn, and based on the optimized best cover image, implement differentiated short video delivery for different user groups.

[0045] The method for sequentially optimizing the image quality of each best cover image includes: Connectivity component labeling methods (such as breadth-first search and depth-first search) are used to perform connectivity analysis on the salient points corresponding to each best cover image, thereby obtaining the salient regions corresponding to each best cover image. It should be noted that the salient points used in the connectivity analysis are all salient points in the set of salient points obtained by using the salient point analysis model. Pixels in each salient region are all salient points, and the salient region represents the hot spot area of ​​visual attention. In each best cover image, the region formed by pixels that are not salient points is marked as the background region. An adaptive histogram equalization method is used to dynamically adjust the contrast of each salient region in each best cover image to enhance local details and optimize visual effects. Edge-preserving filtering algorithms (such as bilateral filtering, weighted median filtering, etc.) are applied to perform global consistency processing on each best cover image to ensure a natural transition between salient and background regions, thereby improving image quality. The adaptive histogram equalization method and edge-preserving filtering algorithm are existing technologies, and the specific process will not be described in detail here.

[0046] This embodiment achieves personalized optimal cover image selection for different user groups by performing frame-by-frame processing, subject detection, and image feature extraction on short videos, combined with multi-dimensional feature analysis of target users. This enhances the targeting and attractiveness of short video cover images. After determining the optimal cover image, saliency analysis and regional optimization design methods are used to automatically locate visual attention areas and rationally arrange cover text, fully satisfying the information preferences of different user groups and effectively improving the overall visual appeal of the short video. Image quality optimization for each optimal cover image involves local detail enhancement and global consistency processing, which not only improves the visual refinement of the cover image but also ensures clear readability in complex backgrounds, thereby improving the conversion rate of short video delivery. It realizes a precise delivery strategy based on content attributes and user characteristics. Compared with traditional methods that rely on behavioral data and tag matching, it can more accurately grasp user needs, thereby significantly improving the click-through rate and user interaction of video content, optimizing the accuracy and effectiveness of short video delivery, and ultimately enhancing the reach and dissemination value of short video content. Example 2:

[0047] This application also provides an electronic device. The electronic device may include one or more processors and one or more memories. The memories store computer-readable code, which, when executed by the one or more processors, can perform an image recognition and content optimization method for short video platform streaming as described above.

[0048] The method or system according to the embodiments of this application can also be implemented using the architecture of the electronic device shown in this application. The electronic device may include a bus, one or more CPUs, ROM, RAM, a communication port connected to a network, input / output, a hard disk, etc. The storage device in the electronic device, such as a ROM or hard disk, may store the image recognition and content optimization method for short video platform streaming provided in this application. Furthermore, the electronic device may also include a user interface. Of course, the architecture shown in this application is merely exemplary; when implementing different devices, one or more components of the electronic device shown in this application may be omitted according to actual needs. Example 3:

[0049] Please refer to the accompanying drawings. One embodiment of this application discloses a computer-readable storage medium. The computer-readable storage medium stores computer-readable instructions. When the computer-readable instructions are executed by a processor, an image recognition and content optimization method for short video platform streaming according to an embodiment of this application, as described above, can be performed. The storage medium includes, but is not limited to, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.

[0050] Furthermore, according to embodiments of this application, the processes described in the above-referenced flowcharts can be implemented as computer software programs. For example, this application provides a non-transitory machine-readable storage medium storing machine-readable instructions that can be executed by a processor to perform instructions corresponding to the method steps provided in this application, such as an image recognition and content optimization method for streaming on a short video platform. When this computer program is executed by a central processing unit (CPU), it performs the functions defined in the method of this application.

[0051] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

[0052] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0053] In the description of this invention, it should be understood that the terms "first," "second," etc., are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance.

[0054] In the description of this invention, unless otherwise stated, "a plurality of" means two or more.

[0055] In the description of this invention, "several" means one or more, and "a large number" means two or more.

[0056] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0057] All formulas in this manual are dimensionless and calculated numerically. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters and thresholds in the formulas are set by those skilled in the art according to the actual situation.

[0058] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.

Claims

1. A method for image recognition and content optimization for streaming on a short video platform, characterized in that, include: Step S1: Perform frame segmentation on the short video and extract the video image sequence; Step S2: Perform subject detection on each frame of the video image sequence in sequence, identify the corresponding subject content, and extract image feature data based on the subject content; Step S3: Perform multi-dimensional feature segmentation on the target users to obtain common feature data corresponding to different user groups; Step S4: Perform feature matching between image feature data and common feature data, and select the best cover image that matches different user groups from the video images based on the matching results; Step S5: Perform a salience analysis on each best cover image in turn, locate the visual attention area, and place the designed cover text in the visual attention area of ​​each best cover image; Step S6: Optimize the image quality of each best cover image in turn, and based on the optimized best cover image, implement differentiated short video delivery for different user groups.

2. The image recognition and content optimization method for short video platform streaming according to claim 1, characterized in that, Methods for identifying the main content corresponding to a video image include: The video image is input into the trained subject detection model, and the subject detection result is output, which includes the subject localization rectangle and the subject type label. The video image is cropped based on the main subject positioning rectangle to obtain the main subject region within the video image; the main subject region is combined with the main subject type tag to form the main subject content. Methods for extracting image feature data based on main content include: Color analysis is performed on the main content to extract color distribution and main color tone; composition and style analysis is performed on the video images to extract visual style tags, subject position, and image balance; the subject type tags, color distribution, main color tone, visual style tags, subject position, and image balance are integrated to form image feature data.

3. The image recognition and content optimization method for short video platform streaming according to claim 2, characterized in that, Methods for extracting color distribution include: Convert the main area from RGB color space to HSV color space, obtain all pixels in the main area, and represent each pixel with the corresponding HSV color value; K-means clustering is used to cluster all pixels to obtain multiple color clusters. The number of pixels in each color cluster is counted and labeled as the pixel count. The sum of all pixel counts is calculated to obtain the total number of pixels. For each color cluster, the mean of the HSV color values ​​of all pixels is calculated to obtain the color center. The ratio of the number of pixels in each cluster to the total number of pixels is used as the color proportion of the corresponding color center. Based on all color centers and their corresponding color proportions, the color distribution of the main region is constructed. Methods for extracting image balance include: The video image is divided into nine equal regions according to the rule of thirds. The corresponding color tone of each region is extracted sequentially, and the brightness value of the corresponding color tone in each region is obtained. Each brightness value is normalized to obtain a standard brightness value. Based on each standard brightness value, the color of each region is calculated. The area of ​​the overlapping portion between each region and the main body region is calculated and marked as the intersection area. The image width and height of the video image are obtained, and the area of ​​each region is calculated. The ratio of each intersection area to the corresponding region area is calculated to obtain the area proportion. The product of the region color, area proportion, and a preset weight for each region is calculated to obtain the visual weight of each region. The coordinates of the center of each equally divided region are calculated based on the image width and image height. The coordinates of the center of all regions are weighted and averaged to obtain the coordinates of the visual centroid. The coordinates of the geometric center of the video image are calculated based on the image width and image height, and the Euclidean distance is used to measure the actual deviation between the coordinates of the visual centroid and the geometric center. The maximum deviation is calculated, and the image balance is calculated based on the actual deviation and the maximum deviation.

4. The image recognition and content optimization method for short video platform streaming according to claim 3, characterized in that, Methods for obtaining common feature data corresponding to different user groups include: Acquire the historical short videos corresponding to each target user, as well as the user behavior data corresponding to each historical short video; among which, the user behavior data includes viewing time and interaction methods, including likes, comments and shares; Obtain the total duration of each historical short video, calculate the ratio of each viewing duration to the corresponding total video duration, and obtain the video completion rate of each historical short video; compare each video completion rate with a preset completion threshold, and delete historical short videos with a completion rate less than the completion threshold. Based on video completion rate and interaction method, calculate the interest weight corresponding to each historical short video; obtain the cover image of each historical short video and mark it as a historical cover image; extract the corresponding image feature data from each historical cover image and mark it as historical feature data; based on the interest weight, calculate the weighted average of the historical feature data corresponding to the historical short videos of the same target user to obtain the interest feature data corresponding to each target user. Based on the interest feature data, the K-means clustering method is used to cluster all target users of the traffic delivery, resulting in multiple user clusters, which are then labeled as user groups. For each user group, the mean of the interest feature data corresponding to all target users of the traffic delivery is calculated to obtain the common feature data corresponding to each user group.

5. The image recognition and content optimization method for short video platform streaming according to claim 4, characterized in that, Methods for calculating the interest weights corresponding to historical short videos include: Sentiment analysis is performed on comments in interactive content to obtain the emotional state of the target users. A weight set is preset, which includes weight coefficients for likes, shares, and different emotional states. The viewing time of historical short videos is obtained, and the difference between the current time and the viewing time is calculated to obtain the time difference. A time decay function is preset, and the time difference is substituted into the time decay function to obtain the time weight. The weight coefficients corresponding to likes, shares, and emotional states are added in sequence to obtain the behavior weight. The behavior weight, video completion rate, and time weight are multiplied in sequence to obtain the interest weight.

6. The image recognition and content optimization method for short video platform streaming according to claim 5, characterized in that, Methods for feature matching between image feature data and common feature data include: Based on all image feature data and common feature data, construct multiple sets of different feature data; for each set of common feature data, calculate the degree of difference between each data point in the image feature data and the corresponding data point in the common feature data. For each type of data in the image feature data, a corresponding data weight is dynamically assigned; based on the data weight, all the degree of difference corresponding to the same feature data set are weighted and summed to obtain the overall matching degree corresponding to each set of feature data.

7. The image recognition and content optimization method for short video platform streaming according to claim 6, characterized in that, Methods for dynamically assigning data weights include: For each user group, the local dispersion of each type of data in the image feature data is calculated sequentially; the mean of the local dispersion of the same data type in the corresponding image feature data is calculated to obtain the average dispersion of each type of data in the image feature data, and the reciprocal of the average dispersion is used as the corresponding data weight.

8. The image recognition and content optimization method for short video platform streaming according to claim 7, characterized in that, Methods for selecting the best cover images from video images to match different user groups include: The overall matching degree corresponding to the same common feature data is compared, and the video image corresponding to the overall matching degree with the highest value is used as the candidate cover image for the corresponding user group; the overall matching degree corresponding to each candidate cover image is compared with the preset matching threshold. If the overall matching degree is greater than or equal to the matching threshold, the corresponding candidate cover image will be used as the best cover image for the corresponding user group. If the overall matching degree is less than the matching threshold, the corresponding user group will be marked as the group to be analyzed; Based on the common feature data of the group to be analyzed, a style transfer model is used to sequentially perform style transfer on all video images to obtain transferred images. The corresponding image feature data is extracted from each transferred image and marked as transferred feature data. Each set of transferred feature data is matched with the common feature data of the group to be analyzed to obtain the overall matching degree corresponding to each transferred image and marked as the transfer matching degree. All transfer matching degrees are compared, and the transferred image corresponding to the transfer matching degree with the largest value is used as the best cover image for the corresponding group to be analyzed.

9. The image recognition and content optimization method for short video platform streaming according to claim 8, characterized in that, Methods for locating visual attention areas include: A pre-trained saliency analysis model is used to perform saliency analysis on each best cover image in turn, obtaining the set of salient points corresponding to each best cover image. The set of salient points includes multiple salient parameters corresponding to the salient points, including position coordinates and saliency intensity. According to the cover text, the region parameters of the corresponding text area are determined. Based on the set of salient points and region parameters corresponding to each best cover image, the text area corresponding to each best cover image is constructed in turn. In each best cover image, the corresponding text area is mapped sequentially. The area of ​​the overlapping part between each text area and the main body area is calculated and marked as the overlapping area. Each overlapping area is compared with a preset area threshold, and salient points with overlapping areas greater than the area threshold are deleted from all salient point sets. The deleted salient point set is marked as the candidate point set. Based on the candidate point set corresponding to each best cover image and the area parameters, the text area corresponding to each best cover image is reconstructed and marked as the reconstructed area. The corresponding reconstructed area is added sequentially to each best cover image, and the cover text is added to each reconstructed area to obtain the text cover image. The image balance of each cover image is extracted sequentially. Based on a preset set of proportions, the image balance of each cover image and the salience intensity of the corresponding salient points are weighted and summed to obtain the layout quality of each cover image. All layout quality of the same cover image are compared, and the reconstructed area with the highest layout quality is taken as the visual attention area of ​​the corresponding cover image.

10. The image recognition and content optimization method for short video platform streaming according to claim 9, characterized in that, The method for sequentially optimizing the image quality of each best cover image includes: The connected component labeling method is used to perform connectivity analysis on the salient points corresponding to each best cover image, thereby obtaining the salient regions corresponding to each best cover image; In each best cover image, the region consisting of pixels that are not salient points is marked as the background region; an adaptive histogram equalization method is used to dynamically adjust the contrast of each salient region in each best cover image; and an edge-preserving filtering algorithm is applied to perform global consistency processing on each best cover image to complete the image quality optimization of each best cover image.

Citation Information

Patent Citations

  • Image generation method and device and storage medium

    CN111415396A

  • Video processing method, storage medium and processor

    CN113382301A

  • Short video immersive advertisement promotion method and system based on computer vision

    CN115131065A

  • Method and device for providing video cover, electronic equipment and computer readable medium

    CN118042186A

  • Multi-factor driven micro map image visual balance degree calculation method

    CN119107318A