Artificial intelligence-based video content processing method and system
By using an AI-based video content processing method and a neural network model to automatically extract video feature information, this method solves the problems of traditional methods being time-consuming and susceptible to subjective factors, and achieves fast and accurate video information extraction and modern information retrieval.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN INSIGMA INNOVATION SOFTWARE CO LTD
- Filing Date
- 2025-01-14
- Publication Date
- 2026-05-19
AI Technical Summary
Traditional video content processing methods rely on manual intervention, are time-consuming, and are easily affected by subjective factors, failing to meet the needs of modern information retrieval.
An AI-based video content processing method is adopted, which collects historical videos for preprocessing, extracts feature vectors and sets labels, uses a neural network model to learn video feature information, and applies it to the current video to automatically extract feature information.
It enables fast and accurate extraction of video information, reduces the time and subjective influence of manual annotation, and meets the needs of modern information retrieval.
Smart Images

Figure CN119992413B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of video processing technology, specifically to a video content processing method and system based on artificial intelligence. Background Technology
[0002] With the continuous emergence of various media formats, the amount of video content has increased dramatically, making cross-media storytelling and content creation increasingly important. This necessitates video editors mastering more skills and tools to effectively create and edit content across different media platforms; therefore, the ability to quickly and accurately extract the necessary information from a large volume of videos has become increasingly crucial.
[0003] Traditional video content processing methods typically involve manual intervention. Professionals watch the full video and manually mark areas of interest, such as specific objects or scenes, before dividing these areas. While this method can yield relatively accurate results, manual annotation is time-consuming and easily affected by subjective factors, thus failing to meet the needs of modern information retrieval.
[0004] Therefore, since it does not meet the existing needs, we propose a video content processing method and system based on artificial intelligence. Summary of the Invention
[0005] The purpose of this invention is to provide a video content processing method and system based on artificial intelligence. This method involves collecting different types of historical videos, extracting feature vectors from the videos using a color space method, and assigning corresponding labels to these feature vectors. It then determines the information expressed by each historical video and organizes the information, feature vectors, and labels of each historical video as input for machine learning. A neural network model learns the feature information of each historical video and outputs the information expressed by each historical video. The learned neural network model is then applied to the current video, enabling it to automatically extract the feature information from the current video and output the information expressed by the current video. This avoids the drawbacks of time-consuming manual annotation and susceptibility to subjective factors, thereby meeting the needs of modern information retrieval and solving the problems mentioned in the background section.
[0006] To achieve the above objectives, the present invention provides the following technical solution:
[0007] The video content processing method based on artificial intelligence includes the following steps:
[0008] Step 1: Collect different types of historical videos and preprocess each video. Preprocessing includes: video editing, compression and annotation, as well as color balancing, filtering, noise reduction, contrast enhancement and artifact removal.
[0009] Step 2: Extract feature information from each historical video using the color space method, and use it as the feature vector for each historical video;
[0010] Step 3: Analyze the content of each historical video and assign corresponding labels to the feature vectors of each historical video; identify the events in each historical video and determine the information expressed by each historical video; and organize the information, feature vectors, and labels expressed by each historical video as the input set for machine learning.
[0011] Step 4: Create a neural network model. Input the input set into the neural network model for training, so that the neural network model learns the feature information of each historical video and outputs the information expressed by each historical video.
[0012] Step 5: Apply the learned neural network model to the current video. The neural network model automatically learns the feature information in the current video and outputs the information expressed in the current video for the user to view.
[0013] Furthermore, step one, which involves preprocessing each historical video, specifically includes the following steps:
[0014] Video editing: Divide each historical video into multiple segments as needed;
[0015] Video compression: Compress the size of each historical video file;
[0016] Annotation processing: Annotate the keyframes in each historical video to record the location information of the events, people, and objects that occurred in that frame;
[0017] Color balance: Eliminates color shift and unevenness in each historical video, ensuring that each historical video presents a consistent color tone;
[0018] Filtering: Gaussian filtering was selected to be applied to each historical video.
[0019] Noise reduction: Mean filtering is used to reduce noise caused by uneven lighting in each historical video.
[0020] Contrast Enhancement: Use contrast enhancement algorithms to enhance the contrast in each historical video, distinguishing objects and scenes in different brightness areas;
[0021] Artifact removal: Removes blurred and artifact areas from each historical video, restoring clear facial and object details from each historical video.
[0022] Furthermore, step two, which involves extracting feature information from each historical video using a color space method, specifically includes the following steps:
[0023] Select the HSV space and convert each frame of each historical video to the HSV space separately.
[0024] Calculate the color distribution of each pixel in each historical video. The color distribution includes: color proportion and color intensity.
[0025] Use a histogram to represent the color proportion of each historical video, and then use the maximum value to represent the color intensity of each historical video;
[0026] The color ratio and color intensity of each historical video are aggregated into a two-dimensional array, which serves as the feature vector for each historical video.
[0027] Furthermore, step three involves analyzing the content of each historical video, specifically including the following steps:
[0028] By automatically classifying the feature vectors of each historical video and assigning corresponding labels to each feature vector;
[0029] By detecting and tracking objects or behaviors in each historical video in real time, intelligent analysis of video content can be achieved.
[0030] By recognizing and translating the speech in each historical video, the video content can be supported in multiple languages;
[0031] By analyzing the behavior in each historical video, we can uncover the information behind the videos;
[0032] By identifying events in each historical video, the timeline of the video content can be processed.
[0033] Furthermore, by recognizing and translating the speech in each historical video, the video content can support multiple languages, including:
[0034] Extract the audio data corresponding to the historical videos;
[0035] Extract the number of language byte syllables appearing in each unit of time corresponding to the speech audio data;
[0036] The time frame coefficient is obtained by using the number of language byte syllables appearing in each unit of time.
[0037] The time frame coefficients are obtained using the following formula:
[0038]
[0039] Where J represents the time frame coefficient; n represents the number of time units contained in the speech audio data, and the value of the time unit ranges from 1 min to 3 min; Ni N represents the number of language byte syllables corresponding to the i-th unit of time; i+1 N represents the number of language byte syllables corresponding to the (i+1)th unit of time; p N represents the average number of language byte syllables corresponding to n units of time; z This represents the intermediate value of the number of language byte syllables corresponding to n units of time.
[0040] The time frame coefficient is compared with a preset time frame coefficient threshold.
[0041] When the time frame coefficient is lower than the preset time frame coefficient threshold, a preset initial time frame length is adopted; and the voice audio data is segmented according to the initial time frame length to obtain audio frame data, and the audio frame data is recognized and translated.
[0042] If the time frame coefficient is not lower than the preset time frame coefficient threshold, then the time frame length is set; and the voice audio data is segmented according to the set time frame length to obtain audio frame data, and the audio frame data is recognized and translated.
[0043] Furthermore, when the time frame coefficient is not lower than a preset time frame coefficient threshold, the time frame length is set, including:
[0044] If the time frame coefficient is not lower than the preset time frame coefficient threshold, then the initial time frame length is retrieved.
[0045] Extract the sound pressure level value corresponding to the audio data of the historical video;
[0046] The initial time frame length is adjusted by combining the sound pressure level corresponding to the voice audio data with the time frame coefficient to obtain the set time frame length.
[0047] The length of the set time frame is obtained by the following formula:
[0048]
[0049] Where T represents the set time frame length; T0 represents the initial time frame length; J represents the time frame coefficient; J y P represents the preset time frame coefficient threshold; xi P represents the sound pressure level at the start of the i-th unit of time; zi P represents the sound pressure level at the end of the i-th unit of time; c This indicates the preset sound pressure reference value.
[0050] Furthermore, step four involves inputting the input set into the neural network model for training, specifically including the following steps:
[0051] The input set is divided into a training set and a test set. The training set is used to train the neural network model, and the test set is used to verify whether the training results of the neural network model are accurate.
[0052] Create a neural network model and initialize the parameters in the network;
[0053] Each sample in the training set is sequentially input into the neural network model, the output value of each neuron is calculated, each output value is compared with the corresponding label, and the loss function is calculated.
[0054] The gradient descent algorithm is used to backpropagate along the gradient of the loss function to update the parameter values of the neural network model.
[0055] After the neural network model has been trained, each sample from the test set is sequentially input into the neural network model to evaluate its performance.
[0056] The forward propagation method inputs all samples from the input set into a neural network model for simulation and prediction, and then uses the neural network model to output the information expressed by each historical video.
[0057] Artificial intelligence-based video content processing systems include:
[0058] The video collection unit is used to collect different types of historical and current videos, and to perform structural adjustments and clarity preprocessing on the collected video content;
[0059] The feature extraction unit is used to extract feature vectors from historical videos and establish corresponding labels; analyze the feature vectors of each historical video to determine the information expressed by each historical video; and then organize the obtained feature information to form the input set of the neural network model.
[0060] The model building unit is used to build a neural network model. The neural network model is trained and tested using the input set to obtain a model that can automatically recognize video content.
[0061] The model practice unit is used to apply the tested neural network model to the current video, automatically identifying the feature content of the current video and the information expressed by the video content.
[0062] Furthermore, it also includes:
[0063] Human-computer interaction terminal, used to receive video content uploaded by users and visually display the processing results of the video content;
[0064] The suggestion feedback unit allows users to provide suggestions and opinions on the video content processing results to the system through a human-computer interaction terminal. If a user disagrees with the current video content processing result, the system will re-identify and process the current video content.
[0065] Furthermore, it also includes:
[0066] The progress indicator unit is used to send progress information to the user after the system outputs the processing results of the video content.
[0067] Compared with the prior art, the beneficial effects of the present invention are:
[0068] This invention collects historical videos of different types, preprocesses each video to ensure clarity and validity, extracts feature vectors from each video using color space, and assigns corresponding labels to these vectors. It then identifies events within each video to determine the information conveyed. The information, feature vectors, and labels from each video are compiled and used as input for machine learning. A neural network model learns the feature information from each video and outputs the information conveyed by each video. This learned neural network model is then applied to the current video, automatically extracting its feature information and outputting the information conveyed by the current video for user viewing. This avoids the drawbacks of time-consuming and subjectively influenced manual annotation, thus meeting the needs of modern information retrieval. Attached Figure Description
[0069] Figure 1 This is a diagram showing the composition of the AI-based video content processing system of the present invention;
[0070] Figure 2 This is a flowchart of the video content processing method based on artificial intelligence according to the present invention. Detailed Implementation
[0071] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0072] To address the technical problem that while existing technologies can yield relatively accurate results, manual annotation is time-consuming and easily influenced by subjective factors, thus failing to meet the needs of modern information retrieval, please refer to [link to relevant documentation]. Figure 1-2 This embodiment provides the following technical solution:
[0073] The video content processing method based on artificial intelligence includes the following steps:
[0074] Step 1: Collect historical videos of different types and preprocess each video. Preprocessing includes video editing, compression, and annotation, as well as color balancing, filtering, noise reduction, contrast enhancement, and artifact removal to improve the accuracy and effectiveness of the analysis. Specifically, this includes the following steps:
[0075] Video Editing: To avoid wasting time and resources, each historical video can be divided into multiple segments using Adobe Premiere Pro software, with the length and size of each segment determined according to actual needs. During editing, avoid including continuously repeating scenes and unnecessary redundant content. Video Compression: Compress each historical video file to reduce its size, saving bandwidth and storage space during storage and transmission. Common compression formats include H.264 and H.265. Annotation: Annotate keyframes in each historical video, recording the events, people, and object positions within each frame, providing a foundation for subsequent video analysis and understanding. Color Balance: Eliminate color shifts and unevenness in each historical video, ensuring a consistent color tone. Specific implementation can be achieved using sRGB or Adobe Premiere Pro. RGB color space offset and ratio adjustment achieves color balance; Filtering: Gaussian filtering is used for each historical video to remove noise and interference, improving the accuracy of analysis and processing; Noise reduction: Mean filtering is used to reduce noise caused by uneven lighting in each historical video, such as flickering caused by traffic, thus removing noise and interference and improving the accuracy of analysis and processing; Contrast enhancement: Contrast enhancement algorithms, such as the Otsu's Threshold algorithm, are used to enhance the contrast in each historical video, distinguishing objects and scenes in different brightness areas, making them easier to analyze and understand; Artifact removal: Blurred and artifact areas are removed from each historical video, such as using the Media Filtering algorithm to restore clear facial and object details in each historical video.
[0076] The beneficial effects achieved by the above are as follows: by preprocessing the video content, the video's feature expression ability is enhanced, and important features in the video content are captured better, thereby improving the accuracy and effectiveness of subsequent analysis and understanding; at the same time, the data is made cleaner, more complete, and more reliable, making the neural network model more stable during the training process, improving the robustness and generalization ability of the neural network model, and thus performing better in practical applications.
[0077] Step 2: Extract feature information from each historical video using the color space method, as the feature vector for each historical video; specifically, this includes the following steps:
[0078] Select the HSV color space and convert each frame of each historical video to the HSV color space separately; specifically, the following formula can be used to convert BGR images to HSV color space:
[0079] HSV[b,g,r]=(0.98*b)+(0.97*g)+(0.62*r);
[0080] Where b, g, and r are the blue, green, and red components in the BGR image, respectively;
[0081] The color distribution of each pixel in each historical video is calculated, including color proportion and color intensity. The color proportion of each historical video is represented by a histogram, and the color intensity of each historical video is represented by the maximum value. Specifically, for each pixel, its color proportion and color intensity in the HSV space are calculated. The color proportion is represented by a histogram, that is, the frequency of occurrence of each color channel divided by the sum of the total frequencies of occurrence. The color intensity can be represented by the maximum value, that is, the maximum value of all color channels is calculated, and the maximum value is taken as the final color intensity.
[0082] The color ratio and color intensity of each historical video are aggregated into a two-dimensional array, which serves as the feature vector for each historical video. The two-dimensional array includes the color ratio matrix and the color intensity matrix of the historical video, to facilitate subsequent video analysis and understanding.
[0083] The beneficial effects achieved by the above are as follows: by utilizing color information, information such as objects, scenes, and actions in the video can be effectively extracted, thereby enabling a better understanding and analysis of the video content; and by using histograms, the distribution and clustering trends of color information in the video can be intuitively displayed, resulting in richer and more comprehensive feature representations, leading to better performance and more accurate results for video content processing.
[0084] Step 3: Analyze the content of each historical video and assign corresponding labels to the feature vectors of each historical video; identify the events in each historical video and determine the information expressed by each historical video; and organize the information, feature vectors, and labels expressed by each historical video as the input set for machine learning; specifically, this includes the following steps:
[0085] By automatically classifying the feature vectors of each historical video and assigning corresponding labels to each feature vector, for example, classifying videos into categories such as pedestrians, vehicles, animals, and buildings based on content, the system facilitates later video analysis and use. Real-time detection and tracking of objects or behaviors in each historical video enables intelligent analysis of video content; for example, real-time detection and labeling of pedestrians facilitates subsequent video analysis and application. Recognition and translation of speech in each historical video provides multilingual support; for example, recognizing and translating Chinese speech into English facilitates cross-cultural communication. Analysis of behavior in each historical video uncovers underlying information; for example, analysis of motion trajectories tracks the direction and speed of moving objects. Event recognition in each historical video allows for timeline processing of the video content; for example, timeline annotation of events in the video enables analysis of the timing and sequence of events.
[0086] The beneficial effects achieved by the above are as follows: by assigning corresponding labels to each feature vector, it can more accurately reflect the video content, making the neural network model easier to understand and interpret; and the labeling of feature vectors helps to improve the interpretability and credibility of the neural network model, thereby improving its reliability and security in practical applications; at the same time, organizing the acquired feature information and related content helps to reduce data redundancy and inconsistency, improve data reusability and utilization, and reduce the training and testing costs of the model.
[0087] Specifically, by recognizing and translating the speech in each historical video, the system enables multilingual support for the video content, including:
[0088] Extract the audio data corresponding to the historical videos;
[0089] Extract the number of language byte syllables appearing in each unit of time corresponding to the speech audio data;
[0090] The time frame coefficient is obtained by using the number of language byte syllables appearing in each unit of time.
[0091] The time frame coefficients are obtained using the following formula:
[0092]
[0093] Where J represents the time frame coefficient; n represents the number of time units contained in the speech audio data, and the value of the time unit ranges from 1 min to 3 min; Ni N represents the number of language byte syllables corresponding to the i-th unit of time; i+1 N represents the number of language byte syllables corresponding to the (i+1)th unit of time; p N represents the average number of language byte syllables corresponding to n units of time; z This represents the intermediate value of the number of language byte syllables corresponding to n units of time.
[0094] The time frame coefficient is compared with a preset time frame coefficient threshold.
[0095] When the time frame coefficient is lower than the preset time frame coefficient threshold, a preset initial time frame length is adopted; and the voice audio data is segmented according to the initial time frame length to obtain audio frame data, and the audio frame data is recognized and translated.
[0096] If the time frame coefficient is not lower than the preset time frame coefficient threshold, then the time frame length is set; and the voice audio data is segmented according to the set time frame length to obtain audio frame data, and the audio frame data is recognized and translated.
[0097] The technical effects of the above solution are as follows: By recognizing and translating speech in historical videos, this solution enables video content to support multiple languages, thereby greatly expanding the audience reach and enhancing the internationalization and multilingual communication capabilities of the video content. The solution calculates time frame coefficients and dynamically adjusts the time frame length based on a comparison of these coefficients with a preset threshold. This dynamic adjustment strategy optimizes segmentation processing based on the actual characteristics of the speech audio data (such as speech rate and syllable density), thus capturing and translating speech content more accurately. By segmenting the speech audio data into more reasonable audio frame data, this solution helps reduce recognition errors and inaccurate translations caused by audio that is too long or too short. The dynamically adjusted time frame length better adapts to the recognition needs of different speech rates and speech characteristics, thereby improving overall recognition and translation accuracy. Multilingual support and accurate speech recognition and translation processing make video content easier to understand and accept, especially for non-native speakers. This helps improve the video viewing experience and user satisfaction. The dynamic adjustment strategy in the solution not only improves processing accuracy but also enhances overall processing efficiency through optimized segmentation processing. At the same time, the solution is highly flexible and can adapt to the video processing needs of different styles and content.
[0098] In summary, this technical solution achieves multilingual support for video content by dynamically adjusting the time frame length and optimizing the segmented processing of audio data, and improves the accuracy and efficiency of speech recognition and translation, thereby enhancing the comprehensibility of video content and user experience.
[0099] Specifically, when the time frame coefficient is not lower than a preset time frame coefficient threshold, the time frame length is set, including:
[0100] If the time frame coefficient is not lower than the preset time frame coefficient threshold, then the initial time frame length is retrieved.
[0101] Extract the sound pressure level value corresponding to the audio data of the historical video;
[0102] The initial time frame length is adjusted by combining the sound pressure level corresponding to the voice audio data with the time frame coefficient to obtain the set time frame length.
[0103] The length of the set time frame is obtained by the following formula:
[0104]
[0105] Where T represents the set time frame length; T0 represents the initial time frame length; J represents the time frame coefficient; J y P represents the preset time frame coefficient threshold; xi P represents the sound pressure level at the start of the i-th unit of time; zi P represents the sound pressure level at the end of the i-th unit of time; c This indicates the preset sound pressure reference value.
[0106] The technical effects of the above solution are as follows: This solution dynamically adjusts the time frame length based on the characteristics of the speech audio data (such as speech rate and language density) by calculating time frame coefficients. This adaptive segmentation processing method helps optimize subsequent speech recognition and translation processes, improving processing accuracy and efficiency. Reasonable audio segmentation ensures that the speech content within each audio frame is relatively independent and complete, which helps reduce errors and ambiguities in the speech recognition process. Simultaneously, segmentation provides clearer speech units for the translation process, contributing to improved accuracy and fluency. Dynamically adjusting the time frame length avoids excessively long audio frames leading to excessive processing time, and also prevents excessively short audio frames from affecting recognition results due to insufficient information. This optimized processing strategy balances processing speed and accuracy, improving overall processing efficiency. By recognizing and translating speech in historical videos, this solution achieves multi-language support for video content. This helps expand the video's audience, enhance its international appeal, and meet the needs of viewers with different language backgrounds. Accurate and fluent speech recognition and translation processing enhances user understanding and acceptance of video content, improving the viewing experience. At the same time, the support for multiple languages makes the video content more inclusive and diverse, which helps to attract more users' attention and participation.
[0107] In summary, this technical solution improves the accuracy and efficiency of speech recognition and translation, optimizes the user experience, and has broad application prospects and value through adaptive audio segmentation processing and multi-language support.
[0108] Step 4: Create a neural network model. Input the input set into the neural network model for training, enabling the model to learn the feature information of each historical video and output the information expressed by each historical video. This specifically includes the following steps:
[0109] The input set is divided into a training set and a test set. The training set is used to train the neural network model, and the test set is used to verify the accuracy of the training results. Simultaneously, data augmentation techniques, such as random cropping, flipping, and rotation, are used to increase the generalization ability of the neural network model and the diversity of the training data. A neural network model is created, and the parameters are initialized to minimize the loss function. Each sample from the training set is sequentially input into the neural network model to learn the feature information and expressive information of each historical video. The output value of each neuron is calculated, and each output value is compared with its corresponding label, and the loss function is calculated. This determines whether the feature information extracted by the neural network model is consistent with the corresponding label. Gradient descent is then applied. The algorithm backpropagates along the gradient of the loss function to update the parameter values of the neural network model. After the neural network model is trained, each sample in the test set is sequentially input into the neural network model to evaluate its performance. Common evaluation metrics include accuracy, precision, recall, and F1 score. If the test results of the neural network model are not ideal, the network structure or hyperparameters, such as the learning rate or batch size, are fine-tuned to achieve better performance. Through forward propagation, all samples in the input set are input into the neural network model for simulated prediction, and the information expressed by each historical video is output by the neural network model. This information can then be used for video retrieval, classification, or generation, thereby achieving more efficient and accurate automated video processing and understanding.
[0110] Step 5: Apply the learned neural network model to the current video. The neural network model automatically learns the feature information in the current video and outputs the information expressed in the current video for the user to view. Specifically, the current video is input into the neural network model for feature information processing, and a result representing the information expressed in the current video is output. The result can be a classification, label, or text description, etc., and then presented to the user to help them better understand the content of the current video. Thus, feature information and expressive information can be automatically extracted directly from the current video without human intervention or supervision. This not only saves time and labor costs but also provides more accurate and reliable processing results, enhancing the user experience.
[0111] The beneficial effects achieved by the above are as follows: By learning the feature information of each historical video through a neural network model, the system can accurately identify objects, scenes, actions, and other information within the video content, enabling in-depth analysis and understanding of the video content. Furthermore, based on the feature information of the video content, it can determine the information conveyed by the video and generate video summaries to help users quickly understand the video content. Alternatively, based on users' interests and viewing history, it can recommend relevant video content, providing users with more intelligent and efficient video processing and understanding tools, and promoting the digital transformation and development of the video industry.
[0112] Artificial intelligence-based video content processing systems include:
[0113] The video collection unit is used to collect different types of historical and current videos, such as manually uploaded, automatically downloaded, and scraped from social media. It performs structural adjustments and clarity preprocessing on the collected video content to ensure data quality and integrity, enabling the neural network model to better extract feature information. This helps reduce model training time and convergence speed, while improving the generalization ability and robustness of the neural network model.
[0114] The feature extraction unit is used to extract feature vectors from historical videos and establish corresponding labels. Specifically, it extracts a set of feature vectors for each historical video to describe the content of the video; classifies the feature vectors and establishes corresponding labels for subsequent training and learning of the neural network model; analyzes the feature vectors of each historical video to determine the information expressed by each historical video; by analyzing the information expressed by each historical video, it can better understand its content and intent, thereby better guiding the training and prediction of the neural network model; and then organizes the obtained feature information to form the input set of the neural network model. The input set includes a list or array of all feature vectors for the neural network model to train and infer.
[0115] The model building unit is used to construct neural network models. Based on the system requirements and input set, it selects a suitable neural network architecture for model building, such as a convolutional neural network model. It then trains and tests the neural network model using the input set to obtain a model capable of automatically recognizing video content. During the training process, it optimizes the model's loss function to improve the model's accuracy and stability, thus obtaining a model capable of automatically recognizing video content.
[0116] The model practice unit is used to apply the tested neural network model to the current video, automatically identifying the feature content and the information expressed in the video. Specifically, the current video is preprocessed to help the neural network model better understand the video content. The neural network model extracts the feature information of the current video, infers and predicts the information to be expressed in the current video based on the feature information, and outputs the final result for users to view and understand.
[0117] The human-computer interaction terminal is used to receive video content uploaded by users and visually display the processing results. Specifically, the terminal provides users with a video preview window to help them view and evaluate the uploaded video content, and offers various control functions such as playback speed adjustment, pause, fast forward, and rewind. It also provides a user feedback mechanism to gather user experience and feedback, including online messaging, rating, and commenting tools to enhance user interaction with the system.
[0118] The feedback unit allows users to provide suggestions and opinions on the video content processing results through a human-computer interaction terminal. If a user disagrees with the processing result, the system will re-process the video content. Specifically, by providing a feedback mechanism, users can participate in the video selection and processing process. For example, the system can provide functions such as search, recommendation, and rating, allowing users to select and use videos according to their needs and preferences. Users can also evaluate the processed video, enabling the system to determine whether the processed video content meets their needs.
[0119] The progress indicator unit is used to send progress information to the user after the system outputs the processing results of the video content, so that the user can check it in a timely manner.
[0120] The beneficial effects achieved by the above content are: by using the above methods, the drawbacks of manual annotation being time-consuming and easily affected by subjective factors are effectively avoided, thereby meeting the needs of modern information retrieval.
[0121] Working principle: By collecting different types of historical videos and preprocessing each video, feature vectors are extracted from each video using color space and labeled accordingly. Events in each video are then identified to determine the information conveyed. The acquired feature information is organized and used as an input set. The neural network model learns the feature information of each video and outputs the information conveyed by each video. The optimized neural network model is then applied to the current video to automatically extract the feature information and output the information conveyed by the current video for user viewing.
[0122] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "possessing," or any other variations thereof are intended to cover non-exclusive possession, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus.
[0123] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A video content processing method based on artificial intelligence, characterized in that, Includes the following steps: Step 1: Collect different types of historical videos and preprocess each video. Preprocessing includes: video editing, compression and annotation, as well as color balancing, filtering, noise reduction, contrast enhancement and artifact removal. Step 2: Extract feature information from each historical video using the color space method, and use it as the feature vector for each historical video; Step 3: Analyze the content of each historical video and assign corresponding labels to the feature vectors of each historical video; identify the events in each historical video and determine the information expressed by each historical video; and organize the information, feature vectors, and labels expressed by each historical video as the input set for machine learning; specifically, this includes the following steps: By automatically classifying the feature vectors of each historical video and assigning corresponding labels to each feature vector; by real-time detection and tracking of objects or behaviors in each historical video, intelligent analysis of video content is achieved; by recognizing and translating the speech in each historical video, multi-language support for video content is achieved; by analyzing the behavior in each historical video, information behind the video is mined; and by identifying events in each historical video, the timeline of video content is processed. This involves recognizing and translating the speech in each historical video to enable multilingual support for the video content, including: Extract the audio data corresponding to the historical video; extract the number of language byte syllables appearing in each unit of time corresponding to the audio data; obtain the time frame coefficient using the number of language byte syllables appearing in each unit of time; wherein, the time frame coefficient is obtained by the following formula: Where J represents the time frame coefficient; n represents the number of time units contained in the speech audio data, and the value of the time unit ranges from 1 min to 3 min; N i N represents the number of language byte syllables corresponding to the i-th unit of time; i+1 N represents the number of language byte syllables corresponding to the (i+1)th unit of time; p N represents the average number of language byte syllables corresponding to n units of time; z This represents the intermediate value of the number of language byte syllables corresponding to n units of time. The time frame coefficient is compared with a preset time frame coefficient threshold; when the time frame coefficient is lower than the preset time frame coefficient threshold, a preset initial time frame length is adopted; and the voice audio data is segmented according to the initial time frame length to obtain audio frame data, and the audio frame data is recognized and translated. When the time frame coefficient is not lower than the preset time frame coefficient threshold, the time frame length is set; and the voice audio data is segmented according to the set time frame length to obtain audio frame data, and the audio frame data is recognized and translated. Step 4: Create a neural network model. Input the input set into the neural network model for training, so that the neural network model learns the feature information of each historical video and outputs the information expressed by each historical video. Step 5: Apply the learned neural network model to the current video. The neural network model automatically learns the feature information in the current video and outputs the information expressed in the current video for the user to view.
2. The video content processing method based on artificial intelligence according to claim 1, characterized in that: Step one involves preprocessing each historical video, specifically including the following steps: Video editing: Divide each historical video into multiple segments as needed; Video compression: Compress the size of each historical video file; Annotation processing: Annotate the keyframes in each historical video to record the location information of the events, people, and objects that occurred in that frame; Color balance: Eliminates color shift and unevenness in each historical video, ensuring that each historical video presents a consistent color tone; Filtering: Gaussian filtering was selected to be applied to each historical video. Noise reduction: Mean filtering is used to reduce noise caused by uneven lighting in each historical video. Contrast Enhancement: Use contrast enhancement algorithms to enhance the contrast in each historical video, distinguishing objects and scenes in different brightness areas; Artifact removal: Removes blurred and artifact areas from each historical video, restoring clear facial and object details from each historical video.
3. The video content processing method based on artificial intelligence according to claim 1, characterized in that: Step two involves extracting feature information from each historical video using a color space method, specifically including the following steps: Select the HSV space and convert each frame of each historical video to the HSV space separately. Calculate the color distribution of each pixel in each historical video. The color distribution includes: color proportion and color intensity. Use a histogram to represent the color proportion of each historical video, and then use the maximum value to represent the color intensity of each historical video; The color ratio and color intensity of each historical video are aggregated into a two-dimensional array, which serves as the feature vector for each historical video.
4. The video content processing method based on artificial intelligence according to claim 1, characterized in that: When the time frame coefficient is not lower than the preset time frame coefficient threshold, the time frame length is set, including: If the time frame coefficient is not lower than the preset time frame coefficient threshold, then the initial time frame length is retrieved. Extract the sound pressure level value corresponding to the audio data of the historical video; The initial time frame length is adjusted by combining the sound pressure level corresponding to the voice audio data with the time frame coefficient to obtain the set time frame length. The length of the set time frame is obtained by the following formula: Where T represents the set time frame length; T0 represents the initial time frame length; J represents the time frame coefficient; J y P represents the preset time frame coefficient threshold; xi P represents the sound pressure level at the start of the i-th unit of time; zi P represents the sound pressure level at the end of the i-th unit of time; c This indicates the preset sound pressure reference value.
5. The video content processing method based on artificial intelligence according to claim 1, characterized in that: Step four involves feeding the input set into the neural network model for training, and specifically includes the following steps: The input set is divided into a training set and a test set. The training set is used to train the neural network model, and the test set is used to verify whether the training results of the neural network model are accurate. Create a neural network model and initialize the parameters in the network; Each sample in the training set is sequentially input into the neural network model, the output value of each neuron is calculated, each output value is compared with the corresponding label, and the loss function is calculated. The gradient descent algorithm is used to backpropagate along the gradient of the loss function to update the parameter values of the neural network model. After the neural network model has been trained, each sample from the test set is sequentially input into the neural network model to evaluate its performance. The forward propagation method inputs all samples from the input set into a neural network model for simulation and prediction, and then uses the neural network model to output the information expressed by each historical video.
6. An artificial intelligence-based video content processing system, applied to the artificial intelligence-based video content processing method according to any one of claims 1-5, characterized in that: include: The video collection unit is used to collect different types of historical and current videos, and to perform structural adjustments and clarity preprocessing on the collected video content; The feature extraction unit is used to extract feature vectors from historical videos and establish corresponding labels; analyze the feature vectors of each historical video to determine the information expressed by each historical video; and then organize the obtained feature information to form the input set of the neural network model. The model building unit is used to build a neural network model. The neural network model is trained and tested using the input set to obtain a model that can automatically recognize video content. The model practice unit is used to apply the tested neural network model to the current video, automatically identifying the feature content of the current video and the information expressed by the video content.
7. The video content processing system based on artificial intelligence according to claim 6, characterized in that: Also includes: Human-computer interaction terminal, used to receive video content uploaded by users and visually display the processing results of the video content; The suggestion feedback unit allows users to provide suggestions and opinions on the video content processing results to the system through a human-computer interaction terminal. If a user disagrees with the current video content processing result, the system will re-identify and process the current video content.
8. The video content processing system based on artificial intelligence according to claim 6, characterized in that: Also includes: The progress indicator unit is used to send progress information to the user after the system outputs the processing results of the video content.