Video content processing method and system based on artificial intelligence
Through an artificial intelligence-based method, the neural network model is used to automatically extract video feature information, which solves the problem that traditional video content processing methods are time-consuming and susceptible to subjective factors, and achieves efficient and accurate video content processing.
Patent Information
- Application Number
- CN202510058159.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-14
AI Technical Summary
Traditional video content processing methods rely on manual intervention, are time-consuming and susceptible to subjective factors, and cannot meet the needs of modern information retrieval.
Using an artificial intelligence-based method, we collect and preprocess historical videos, extract feature vectors using the color space method, set labels for them, sort them into input sets for machine learning, and train neural network models to automatically extract video feature information.
It realizes automated video content processing, avoids the time and subjective influence of manual labeling, and meets the needs of modern information retrieval.
Smart Images

Figure CN119992413A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video processing technology, and in particular to a video content processing method and system based on artificial intelligence. Background Art
[0002] With the continuous emergence of various media forms, the amount of video content has increased dramatically, and cross-media narrative and content production have become increasingly important. Video editors need to master more skills and tools to effectively produce and edit content on different media platforms; therefore, how to quickly and accurately find the required information from a large number of videos has become increasingly important.
[0003] Traditional video content processing methods usually involve manual intervention. Professionals watch the full video and manually mark out areas of interest, such as specific objects and scenes, and then divide these areas. Although this method can produce more accurate results, manual labeling is time-consuming and easily affected by personal subjective factors, and thus cannot meet the needs of modern information retrieval.
[0004] Therefore, the existing needs are not met, and we propose a video content processing method and system based on artificial intelligence. Summary of the invention
[0005] The purpose of the present invention is to provide a video content processing method and system based on artificial intelligence, by collecting different types of historical videos, using the color space method to extract the feature vector of the video, and setting corresponding labels for the feature vector; judging the information expressed by each historical video, and organizing the information, feature vector and label expressed by each historical video as an input set for machine learning; making the neural network model learn the feature information of each historical video, and output the information expressed by each historical video; then applying the learned neural network model to the current video, so that it automatically extracts the feature information in the current video, and outputs the information expressed by the current video; thereby avoiding the disadvantages of manual labeling that is time-consuming and easily affected by subjective factors, thereby meeting the needs of modern information retrieval and solving the problems raised in the above-mentioned background technology.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] The video content processing method based on artificial intelligence includes the following steps:
[0008] Step 1: Collect different types of historical videos and pre-process each historical video. The pre-processing includes: video editing, compression and labeling, as well as color balancing, filtering, noise reduction, contrast enhancement and artifact removal.
[0009] Step 2: Using the color space method to extract feature information from each historical video as a feature vector of each historical video;
[0010] Step 3: Analyze the content of each historical video and set corresponding labels for the feature vectors of each historical video; identify the events in each historical video and determine the information expressed by each historical video; and organize the information, feature vectors and labels expressed by each historical video as the input set for machine learning;
[0011] Step 4: Create a neural network model, input the input set into the neural network model for training, so that the neural network model learns the feature information of each historical video and outputs the information expressed by each historical video;
[0012] Step 5: Apply the learned neural network model to the current video, automatically learn the feature information in the current video through the neural network model, and output the information expressed by the current video for users to view.
[0013] Furthermore, the step 1 is to pre-process each historical video, which specifically includes the following steps:
[0014] Video editing: split each historical video into multiple clips according to needs;
[0015] Video compression: compress the size of each historical video file;
[0016] Annotation processing: Annotate the key frames in each historical video and record the events, people, and location information of objects in the frame;
[0017] Color balance: Eliminate color drift and unevenness of each historical video, so that each historical video presents a consistent color tone;
[0018] Filtering: Choose to use Gaussian filtering to filter each historical video;
[0019] Noise reduction: The mean filter method is used to reduce the noise caused by uneven light in each historical video;
[0020] Contrast enhancement: Use contrast enhancement algorithms to enhance the contrast in each historical video and distinguish objects and scenes in different brightness areas;
[0021] Artifact removal: Remove blurry and artifact areas in each historical video, and restore clear facial and object details in each historical video.
[0022] Furthermore, the step 2 uses a color space method to extract feature information from each historical video, specifically including the following steps:
[0023] Select HSV space and convert each frame image of each historical video into HSV space;
[0024] Calculate the color distribution of each pixel in each historical video. The color distribution includes: color ratio and color intensity;
[0025] Use a histogram to represent the color ratio of each historical video, and then use the maximum value to represent the color intensity of each historical video;
[0026] The color ratio and color intensity of each historical video are summarized into a two-dimensional array as the feature vector of each historical video.
[0027] Furthermore, the step three, analyzing the content of each historical video, specifically includes the following steps:
[0028] By automatically classifying the feature vectors of each historical video and setting corresponding labels for each feature vector;
[0029] By real-time detection and tracking of objects or behaviors in each historical video, intelligent analysis of video content is achieved;
[0030] By recognizing and translating the voice in each historical video, the video content can be supported in multiple languages;
[0031] By analyzing the behaviors in each historical video, we can mine the information behind the video;
[0032] By identifying events in each historical video, the timeline of the video content can be processed.
[0033] Furthermore, by recognizing and translating the speech in each historical video, we can support multiple languages for the video content, including:
[0034] Extracting the voice and audio data corresponding to the historical video;
[0035] Extracting the number of language byte syllables appearing per unit time corresponding to the speech audio data;
[0036] Obtaining a time frame coefficient using the number of language byte syllables appearing in each unit time;
[0037] The time frame coefficient is obtained by the following formula:
[0038]
[0039] Wherein, J represents the time frame coefficient; n represents the number of unit time contained in the speech audio data, and the value range of the unit time is 1min-3min; Ni Indicates the number of language byte syllables corresponding to the i-th unit time; N i+1 Indicates the number of language byte syllables corresponding to the i+1th unit time; N p N represents the average number of language syllables corresponding to n units of time; z Indicates the median number of language byte syllables corresponding to n unit time;
[0040] comparing the time frame coefficient with a preset time frame coefficient threshold;
[0041] When the time frame coefficient is lower than a preset time frame coefficient threshold, a preset initial time frame length is adopted; and the speech audio data is segmented according to the initial time frame length, audio frame data is obtained, and the audio frame data is recognized and translated;
[0042] When the time frame coefficient is not lower than a preset time frame coefficient threshold, the time frame length is set; the speech audio data is segmented according to the set time frame length, audio frame data is obtained, and the audio frame data is recognized and translated.
[0043] Further, when the time frame coefficient is not lower than a preset time frame coefficient threshold, setting the time frame length includes:
[0044] When the time frame coefficient is not lower than a preset time frame coefficient threshold, retrieving the initial time frame length;
[0045] Extracting the sound pressure value corresponding to the voice audio data corresponding to the historical video;
[0046] The initial time frame length is adjusted using the sound pressure value corresponding to the speech audio data in combination with a time frame coefficient to obtain a set time frame length;
[0047] The length of the time frame after setting is obtained by the following formula:
[0048]
[0049] Where T represents the time frame length after setting; T0 represents the initial time frame length; J represents the time frame coefficient; J y Indicates the preset time frame coefficient threshold; P xi represents the sound pressure value corresponding to the starting time of the i-th unit time; P zi represents the sound pressure value corresponding to the end time of the i-th unit time; P c Indicates the preset sound pressure reference value.
[0050] Furthermore, the step 4, inputting the input set into the neural network model for training, specifically includes the following steps:
[0051] The input set is divided into a training set and a test set. The training set is used to train the neural network model, and the test set is used to verify whether the training results of the neural network model are accurate.
[0052] Create a neural network model and initialize the parameters in the network;
[0053] Input each sample of the training set into the neural network model in turn, calculate the output value of each neuron, compare each output value with the corresponding label, and calculate the loss function;
[0054] Use the gradient descent algorithm to back-propagate along the gradient of the loss function to update the parameter values of the neural network model;
[0055] After the neural network model training is completed, each sample in the test set is input into the neural network model in turn to evaluate the performance of the neural network model;
[0056] All samples of the input set are input into the neural network model through forward propagation for simulated prediction, and the neural network model is used to output the information expressed by each historical video.
[0057] Artificial intelligence-based video content processing system, including:
[0058] A video collection unit is used to collect different types of historical videos and current videos, and to perform structural adjustment and clarity preprocessing on the collected video content;
[0059] The feature extraction unit is used to extract feature vectors from historical videos and establish corresponding labels; analyze the feature vectors of each historical video to determine the information expressed by each historical video; and then organize the obtained feature information to form an input set of the neural network model;
[0060] A model building unit, used to build a neural network model, train and test the neural network model using an input set, and obtain a model that can automatically recognize video content;
[0061] The model practice unit is used to apply the tested neural network model to the current video to automatically identify the feature content of the current video and the information expressed by the video content.
[0062] Furthermore, it also includes:
[0063] A human-computer interaction terminal is used to receive video content uploaded by users and visualize the processing results of the video content;
[0064] Suggestion feedback unit: users use the human-computer interaction terminal as a carrier to provide the system with suggestions and opinions on the video content processing results; if the user has objections to the processing results of the current video content, the system will re-identify and process the current video content.
[0065] Furthermore, it also includes:
[0066] The progress prompt unit is used to send a progress prompt message to the user after the system outputs the processing result of the video content.
[0067] Compared with the prior art, the present invention has the following beneficial effects:
[0068] The present invention collects different types of historical videos and pre-processes each historical video to ensure the clarity and effectiveness of the video; uses a color space method to extract feature vectors in each historical video and sets corresponding labels for the feature vectors; then identifies events in each historical video and determines the information expressed by the video; and organizes the information, feature vectors and labels expressed by each historical video as an input set for machine learning; enables a neural network model to learn the feature information of each historical video and outputs the information expressed by each historical video; then applies the learned neural network model to a current video so that the model automatically extracts the feature information in the current video and outputs the information expressed by the current video for viewing by a user; thereby avoiding the disadvantages of manual labeling that is time-consuming and easily affected by subjective factors, thereby meeting the needs of modern information retrieval. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 This is a composition diagram of the video content processing system based on artificial intelligence of the present invention;
[0070] Figure 2 The figure is a flow chart of the video content processing method based on artificial intelligence of the present invention. DETAILED DESCRIPTION
[0071] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0072] In order to solve the technical problem that although the existing technology can obtain more accurate results, manual annotation takes a long time and is easily affected by personal subjective factors, thus failing to meet the needs of modern information retrieval, please refer to Figure 1-2 , this embodiment provides the following technical solutions:
[0073] The video content processing method based on artificial intelligence includes the following steps:
[0074] Step 1: Collect different types of historical videos and preprocess each historical video. The preprocessing includes: video editing, compression and annotation processing, as well as color balancing, filtering, noise reduction, contrast enhancement and artifact removal processing to improve the accuracy and effect of the analysis process; the specific steps include:
[0075] Video editing: To avoid wasting time and resources, you can use Adobe Premiere Pro software to split each historical video into multiple clips as needed. The length and size of each clip can be determined according to actual conditions, and when editing, avoid editing consecutive repeated scenes and unnecessary redundant content into the video; Video compression: Compress the size of each historical video file to reduce the size of the video, so as to save bandwidth and storage space during storage and transmission. Common video compression formats include H.264 and H.265; Annotation processing: Annotate the key frames in each historical video to record the events, people, and object location information of the frame, providing a basis for subsequent video analysis and understanding; Color balance: Eliminate color drift and unevenness in each historical video, so that each historical video presents a consistent color tonality; In specific implementation, sRGB or Adobe RGB adjusts the offset and ratio of the color space to achieve color balance; filtering processing: choose to use Gaussian filtering to filter each historical video to remove noise and interference in the video, thereby improving the accuracy of analysis and processing; noise reduction processing: use the mean filtering method to reduce the noise caused by uneven light in each historical video, such as the flicker caused by heavy traffic, thereby removing noise and interference in the video, thereby improving the accuracy of analysis and processing; contrast enhancement: use contrast enhancement algorithms, such as the Otsu_threshold algorithm, to enhance the contrast in each historical video, distinguish objects and scenes in different brightness areas, and make them easier to analyze and understand; artifact removal processing: remove blurred areas and artifact areas in each historical video, such as using the Median Filtering artifact removal algorithm to restore clear facial and object details in each historical video.
[0076] The beneficial effects achieved by the above content are: by preprocessing the video content, the video feature expression ability is made stronger, and the important features in the video content are better captured, thereby improving the accuracy and effect of subsequent analysis and understanding; at the same time, the data is made cleaner, more complete, and more reliable, making the neural network model more stable during the training process, improving the robustness and generalization ability of the neural network model, so that it performs better in practical applications.
[0077] Step 2: Using the color space method to extract feature information from each historical video as a feature vector of each historical video; specifically, the following steps are included:
[0078] Select HSV space and convert each frame of each historical video into HSV space. Specifically, the following formula can be used to convert BGR image into HSV space:
[0079] HSV[b,g,r]=(0.98*b)+(0.97*g)+(0.62*r);
[0080] Among them, b, g, and r are the blue component, green component, and red component in the BGR image respectively;
[0081] Calculate the color distribution of each pixel in each historical video, the color distribution includes: color ratio and color intensity; use a histogram to represent the color ratio of each historical video, and then use the maximum value to represent the color intensity of each historical video; specifically, for each pixel, calculate its color ratio and color intensity in the HSV space; the color ratio is represented by a histogram, that is, the frequency of occurrence of each color channel divided by the sum of the total frequency of occurrence; the color intensity can be represented by the maximum value, that is, the maximum values of all color channels are obtained, and then the maximum value is taken as the final color intensity.
[0082] The color ratio and color intensity of each historical video are summarized into a two-dimensional array as the feature vector of each historical video; wherein the two-dimensional array includes: a color ratio matrix of the historical video and a color intensity matrix of the historical video, so as to facilitate subsequent video analysis and understanding.
[0083] The beneficial effects achieved by the above content are as follows: by utilizing color information, information such as objects, scenes, and actions in the video can be effectively extracted, so as to better understand and analyze the video content; and through histogram representation, the distribution and aggregation trend of color information in the video can be intuitively displayed, so as to obtain a richer and more comprehensive feature representation, bringing better performance and more accurate results to video content processing.
[0084] Step 3: Analyze the content of each historical video and set corresponding labels for the feature vectors of each historical video; identify the events in each historical video and determine the information expressed by each historical video; and organize the information, feature vectors and labels expressed by each historical video as the input set for machine learning; specifically, the following steps are included:
[0085] By automatically classifying the feature vectors of each historical video and setting corresponding labels for each feature vector; for example, based on the video content, it can be divided into multiple categories such as pedestrians, vehicles, animals and buildings, so as to facilitate later video analysis and use; by real-time detection and tracking of objects or behaviors in each historical video, intelligent analysis of video content is achieved; for example, pedestrians in each historical video can be detected in real time and marked when detected, so as to facilitate subsequent video analysis and application; by recognizing and translating the voice in each historical video, multiple language support for video content is achieved; for example, Chinese voice in the video can be recognized and translated into English for cross-cultural communication and use; by analyzing the behavior in each historical video, the information behind the video can be mined; for example, the motion trajectory in the video can be analyzed to track the moving direction and speed of the moving object; by identifying the events in each historical video, the timeline of the video content can be processed; for example, the events in the video can be marked on the timeline to analyze the time and sequence of the events.
[0086] The beneficial effects achieved by the above content are: by setting corresponding labels for each feature vector, it can more accurately reflect the video content, making the neural network model easier to understand and interpret; and setting labeled feature vectors helps to improve the interpretability and credibility of the neural network model, thereby improving its reliability and security in practical applications; at the same time, organizing the acquired feature information and related content helps to reduce data redundancy and inconsistency, improve data reusability and utilization, and reduce the training and testing costs of the model.
[0087] Specifically, by recognizing and translating the speech in each historical video, we can support multiple languages for the video content, including:
[0088] Extracting the voice and audio data corresponding to the historical video;
[0089] Extracting the number of language byte syllables appearing per unit time corresponding to the speech audio data;
[0090] Obtaining a time frame coefficient using the number of language byte syllables appearing in each unit time;
[0091] The time frame coefficient is obtained by the following formula:
[0092]
[0093] Wherein, J represents the time frame coefficient; n represents the number of unit time contained in the speech audio data, and the value range of the unit time is 1min-3min; Ni Indicates the number of language byte syllables corresponding to the i-th unit time; N i+1 Indicates the number of language byte syllables corresponding to the i+1th unit time; N p N represents the average number of language syllables corresponding to n units of time; z Indicates the median number of language byte syllables corresponding to n unit time;
[0094] comparing the time frame coefficient with a preset time frame coefficient threshold;
[0095] When the time frame coefficient is lower than a preset time frame coefficient threshold, a preset initial time frame length is adopted; and the speech audio data is segmented according to the initial time frame length, audio frame data is obtained, and the audio frame data is recognized and translated;
[0096] When the time frame coefficient is not lower than a preset time frame coefficient threshold, the time frame length is set; the speech audio data is segmented according to the set time frame length, audio frame data is obtained, and the audio frame data is recognized and translated.
[0097] The technical effect of the above technical solution is: by recognizing and translating the voice in the historical video, the solution enables the video content to support multiple languages, thereby greatly broadening the audience range of the video and improving the internationalization of the video content and the ability of multilingual communication. In the solution, the time frame coefficient is calculated, and the time frame length is dynamically adjusted according to the comparison result between the coefficient and the preset threshold. This dynamic adjustment strategy can optimize the segmentation processing according to the actual characteristics of the voice audio data (such as speech rate, syllable density, etc.), so as to more accurately capture and translate the voice content. By segmenting the voice audio data into more reasonable audio frame data, the solution helps to reduce the recognition errors and inaccurate translation problems caused by too long or too short audio. The dynamically adjusted time frame length can better adapt to the recognition requirements of different speech rates and voice features, thereby improving the overall recognition and translation accuracy. Multi-language support and accurate speech recognition and translation processing make the video content easier to understand and accept, especially for non-native audiences. This helps to improve the viewing experience and user satisfaction of the video. The dynamic adjustment strategy in the solution not only improves the accuracy of the processing, but also improves the overall processing efficiency by optimizing the segmentation processing. At the same time, the solution has good flexibility and can adapt to video processing requirements of different styles and contents.
[0098] In summary, this technical solution achieves multi-language support for video content by dynamically adjusting the time frame length and optimizing the segmented processing of voice and audio data, and improves the accuracy and efficiency of voice recognition and translation, thereby enhancing the comprehensibility of video content and user experience.
[0099] Specifically, when the time frame coefficient is not lower than a preset time frame coefficient threshold, setting the time frame length includes:
[0100] When the time frame coefficient is not lower than a preset time frame coefficient threshold, retrieving the initial time frame length;
[0101] Extracting the sound pressure value corresponding to the voice audio data corresponding to the historical video;
[0102] The initial time frame length is adjusted using the sound pressure value corresponding to the speech audio data in combination with a time frame coefficient to obtain a set time frame length;
[0103] The length of the time frame after setting is obtained by the following formula:
[0104]
[0105] Where T represents the time frame length after setting; T0 represents the initial time frame length; J represents the time frame coefficient; J y Indicates the preset time frame coefficient threshold; P xi represents the sound pressure value corresponding to the starting time of the i-th unit time; P zi represents the sound pressure value corresponding to the end time of the i-th unit time; P c Indicates the preset sound pressure reference value.
[0106] The technical effect of the above technical solution is: the solution dynamically adjusts the time frame length according to the characteristics of the speech audio data (such as speech speed, language density, etc.) by calculating the time frame coefficient. This adaptive segmentation processing method helps to optimize the subsequent speech recognition and translation process and improve the accuracy and efficiency of the processing. Through reasonable audio segmentation, it can ensure that the speech content in each audio frame is relatively independent and complete, which helps to reduce errors and ambiguities in the speech recognition process. At the same time, segmentation processing also provides clearer speech units for the translation process, which helps to improve the accuracy and fluency of the translation. Dynamic adjustment of the time frame length can avoid long processing time caused by too long audio frames, and can also prevent short audio frames from affecting the recognition effect due to insufficient information. This optimization processing strategy can balance processing speed and accuracy and improve overall processing efficiency. The solution realizes multi-language support for video content by recognizing and translating speech in historical videos. This helps to expand the audience range of videos, improve the internationalization level of videos, and meet the needs of audiences with different language backgrounds. Accurate and smooth speech recognition and translation processing can enhance users' understanding and acceptance of video content and improve the viewing experience. At the same time, multi-language support also makes video content more inclusive and diverse, helping to attract more user attention and participation.
[0107] In summary, this technical solution improves the accuracy and efficiency of speech recognition and translation through adaptive audio segmentation processing and multiple language support, optimizes the user experience, and has broad application prospects and value.
[0108] Step 4: Create a neural network model, input the input set into the neural network model for training, so that the neural network model learns the feature information of each historical video and outputs the information expressed by each historical video; specifically, the following steps are included:
[0109] The input set is divided into a training set and a test set. The training set is used to train the neural network model, and the test set is used to verify whether the training results of the neural network model are accurate. At the same time, data enhancement techniques are used, such as random cropping, flipping, and rotation, to increase the generalization ability of the neural network model and the diversity of training data. A neural network model is created, and the parameters in the network are initialized to minimize the loss function. Each sample of the training set is input into the neural network model in turn to learn the feature information and expression information of each historical video, calculate the output value of each neuron, compare each output value with the corresponding label, and calculate the loss function. In this way, it is determined whether the feature information extracted by the neural network model is consistent with the corresponding label. Then gradient descent is used The algorithm back-propagates along the gradient of the loss function to update the parameter values of the neural network model. After the training of the neural network model is completed, each sample of the test set is input into the neural network model in turn to evaluate the performance of the neural network model. Common evaluation indicators include accuracy, precision, recall, and F1 score. If the test results of the neural network model are not ideal, try to fine-tune the network structure or adjust hyperparameters, such as learning rate or batch size, to obtain better performance. Through forward propagation, all samples of the input set are input into the neural network model for simulated prediction, and the neural network model is used to output the information expressed by each historical video, which can be used in video retrieval, classification or generation, so as to achieve more efficient and accurate automated video processing and understanding.
[0110] Step 5. Apply the learned neural network model to the current video, automatically learn the feature information in the current video through the neural network model, and output the information expressed by the current video for the user to view; specifically, input the current video into the neural network model to process the feature information, and output a result representing the information expressed by the current video; the result content can be a classification, label or text description, etc., and then present the result to the user to help the user better understand the content of the current video; thus, the feature information and expression information can be automatically extracted directly from the current video without manual intervention or supervision; it can not only save time and labor costs, but also provide more accurate and reliable processing results, thereby enhancing the user experience.
[0111] The beneficial effects achieved by the above content are: by learning the characteristic information of each historical video through the neural network model, the objects, scenes, actions and other information in the video content can be accurately identified, and the in-depth analysis and understanding of the video content can be achieved. The information conveyed by the video is obtained based on the characteristic information of the video content, and a video summary is generated to help users quickly understand the video content, or recommend related video content based on the user's interest points and historical viewing records, providing users with more intelligent and efficient video processing and understanding tools, and promoting the digital transformation and development of the video industry.
[0112] Artificial intelligence-based video content processing system, including:
[0113] The video collection unit is used to collect different types of historical and current videos, such as manual uploads, automatic downloads, social media crawls, etc., and to perform structural adjustments and clarity preprocessing on the collected video content to ensure the quality and integrity of the data, so that the neural network model can better extract feature information from it; it helps to reduce model training time and convergence speed, and at the same time improve the generalization ability and robustness of the neural network model.
[0114] The feature extraction unit is used to extract feature vectors from historical videos and establish corresponding labels; specifically, a set of feature vectors is extracted from each historical video to describe the content of the video; the feature vectors are classified and corresponding labels are established to facilitate subsequent training and learning of the neural network model; the feature vectors of each historical video are analyzed to determine the information expressed by each historical video; the information expressed by each historical video is analyzed to better understand its content and intention, thereby better guiding the training and prediction of the neural network model; the obtained feature information is then sorted to form an input set of the neural network model; the input set includes a list or array of all feature vectors to facilitate training and reasoning of the neural network model.
[0115] The model building unit is used to build a neural network model. According to the system requirements and input sets, a suitable neural network architecture is selected to build the model, such as a convolutional neural network model; the neural network model is trained and tested using the input set, so as to obtain a model that can automatically recognize the content of the video; and during the training process, the loss function of the model is optimized to improve the accuracy and stability of the model, so as to obtain a model that can automatically recognize the content of the video.
[0116] The model practice unit is used to apply the tested neural network model to the current video, automatically identify the feature content of the current video and the information expressed by the video content; specifically, preprocess the current video to enable the neural network model to better understand the video content; use the neural network model to extract the feature information of the current video, infer and predict the information to be expressed by the current video based on the feature information, and output the final result for users to view and understand.
[0117] The human-computer interaction terminal is used to receive video content uploaded by users and visualize the processing results of the video content; specifically, the human-computer interaction terminal is used to provide users with a video preview window to help users view and evaluate uploaded video content, and provide multiple speed playback, pause, fast forward, slow rewind and other control functions; at the same time, it also provides a user feedback mechanism to obtain user experience and feedback, specifically: provide online message, rating, comment and other interactive tools to enhance the interactivity between users and the system.
[0118] Suggestion feedback unit, users use the human-computer interaction terminal as a carrier to feedback suggestions and opinions on the video content processing results to the system; if the user has objections to the processing results of the current video content, the system will re-identify and process the current video content; specifically, by providing a feedback mechanism, users can participate in the video selection and processing process; for example: it can provide users with search, recommendation, rating and other functions, allowing users to select and use videos according to their own needs and preferences; it can also evaluate the currently processed video, so that the system can know whether the processed video content meets user needs.
[0119] The progress prompt unit is used to send progress prompt information to the user after the system outputs the processing result of the video content, so that the user can view it in time.
[0120] The beneficial effects achieved by the above content: Through the above method, the disadvantages of manual labeling, which is time-consuming and easily affected by subjective factors, can be effectively avoided, thereby meeting the needs of modern information retrieval.
[0121] Working principle: By collecting different types of historical videos and preprocessing each historical video; using the color space method to extract the feature vector in each historical video, and setting corresponding labels for the feature vectors; then identifying the events in each historical video and judging the information expressed by the video; and organizing the acquired feature information as the input set; enabling the neural network model to learn the feature information of each historical video and output the information expressed by each historical video; and then applying the optimized neural network model to the current video, so that it automatically extracts the feature information in the current video and outputs the information expressed by the current video for users to view.
[0122] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "including", "having" or any other variations thereof are intended to cover non-exclusive possessing, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.
[0123] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A video content processing method based on artificial intelligence, characterized in that: The following steps are involved: Step 1: Collect different types of historical videos and pre-process each historical video. The pre-processing includes: video editing, compression and labeling, as well as color balancing, filtering, noise reduction, contrast enhancement and artifact removal. Step 2: Using the color space method to extract feature information from each historical video as a feature vector of each historical video; Step 3: Analyze the content of each historical video and set corresponding labels for the feature vectors of each historical video; identify the events in each historical video and determine the information expressed by each historical video; and organize the information, feature vectors and labels expressed by each historical video as the input set for machine learning; Step 4: Create a neural network model, input the input set into the neural network model for training, so that the neural network model learns the feature information of each historical video and outputs the information expressed by each historical video; Step 5: Apply the learned neural network model to the current video, automatically learn the feature information in the current video through the neural network model, and output the information expressed by the current video for users to view.
2. The video content processing method based on artificial intelligence according to claim 1, characterized in that: The step 1 is to pre-process each historical video, which specifically includes the following steps: Video editing: split each historical video into multiple clips according to needs; Video compression: compress the size of each historical video file; Annotation processing: Annotate the key frames in each historical video and record the events, people, and location information of objects in the frame; Color balance: Eliminate color drift and unevenness of each historical video, so that each historical video presents a consistent color tone; Filtering: Choose to use Gaussian filtering to filter each historical video; Noise reduction: The mean filter method is used to reduce the noise caused by uneven light in each historical video; Contrast enhancement: Use contrast enhancement algorithms to enhance the contrast in each historical video and distinguish objects and scenes in different brightness areas; Artifact removal: Remove blurry and artifact areas in each historical video, and restore clear facial and object details in each historical video.
3. The video content processing method based on artificial intelligence according to claim 1, characterized in that: The second step is to extract feature information from each historical video using a color space method, which specifically includes the following steps: Select HSV space and convert each frame image of each historical video into HSV space; Calculate the color distribution of each pixel in each historical video. The color distribution includes: color ratio and color intensity; Use a histogram to represent the color ratio of each historical video, and then use the maximum value to represent the color intensity of each historical video; The color ratio and color intensity of each historical video are summarized into a two-dimensional array as the feature vector of each historical video.
4. The video content processing method based on artificial intelligence according to claim 1, characterized in that: The step three is to analyze the content of each historical video, which specifically includes the following steps: By automatically classifying the feature vectors of each historical video and setting corresponding labels for each feature vector; By real-time detection and tracking of objects or behaviors in each historical video, intelligent analysis of video content is achieved; By recognizing and translating the voice in each historical video, the video content can be supported in multiple languages; By analyzing the behaviors in each historical video, we can mine the information behind the video; By identifying events in each historical video, the timeline of the video content can be processed.
5. The video content processing method based on artificial intelligence according to claim 4 is characterized in that: By recognizing and translating the speech in each historical video, the video content can be supported in multiple languages, including: Extracting the voice and audio data corresponding to the historical video; Extracting the number of language byte syllables appearing per unit time corresponding to the speech audio data; Obtaining a time frame coefficient using the number of language byte syllables appearing in each unit time; The time frame coefficient is obtained by the following formula: Wherein, J represents the time frame coefficient; n represents the number of unit time contained in the speech audio data, and the value range of the unit time is 1min-3min; N i Indicates the number of language byte syllables corresponding to the i-th unit time; N i+1 Indicates the number of language byte syllables corresponding to the i+1th unit time; N p N represents the average number of language syllables corresponding to n units of time; z Indicates the median number of language byte syllables corresponding to n unit time; comparing the time frame coefficient with a preset time frame coefficient threshold; When the time frame coefficient is lower than a preset time frame coefficient threshold, a preset initial time frame length is adopted; and the speech audio data is segmented according to the initial time frame length, audio frame data is obtained, and the audio frame data is recognized and translated; When the time frame coefficient is not lower than a preset time frame coefficient threshold, the time frame length is set; the speech audio data is segmented according to the set time frame length, audio frame data is obtained, and the audio frame data is recognized and translated.
6. The video content processing method based on artificial intelligence according to claim 5 is characterized in that: When the time frame coefficient is not lower than a preset time frame coefficient threshold, setting the time frame length includes: When the time frame coefficient is not lower than a preset time frame coefficient threshold, retrieving the initial time frame length; Extracting the sound pressure value corresponding to the voice audio data corresponding to the historical video; The initial time frame length is adjusted using the sound pressure value corresponding to the speech audio data in combination with a time frame coefficient to obtain a set time frame length; The length of the time frame after setting is obtained by the following formula: Where T represents the time frame length after setting; T0 represents the initial time frame length; J represents the time frame coefficient; J y Represents the preset time frame coefficient threshold; P xi represents the sound pressure value corresponding to the starting time of the i-th unit time; P zi represents the sound pressure value corresponding to the end time of the i-th unit time; P c Indicates the preset sound pressure reference value.
7. The video content processing method based on artificial intelligence according to claim 1, characterized in that: The fourth step is to input the input set into the neural network model for training, which specifically includes the following steps: The input set is divided into a training set and a test set. The training set is used to train the neural network model, and the test set is used to verify whether the training results of the neural network model are accurate. Create a neural network model and initialize the parameters in the network; Input each sample of the training set into the neural network model in turn, calculate the output value of each neuron, compare each output value with the corresponding label, and calculate the loss function; Use the gradient descent algorithm to back-propagate along the gradient of the loss function to update the parameter values of the neural network model; After the neural network model training is completed, each sample in the test set is input into the neural network model in turn to evaluate the performance of the neural network model; All samples of the input set are input into the neural network model through forward propagation for simulated prediction, and the neural network model is used to output the information expressed by each historical video.
8. An artificial intelligence-based video content processing system, applied to the artificial intelligence-based video content processing method according to any one of claims 1 to 7, characterized in that: include: A video collection unit is used to collect different types of historical videos and current videos, and to perform structural adjustment and clarity preprocessing on the collected video content; The feature extraction unit is used to extract feature vectors from historical videos and establish corresponding labels; analyze the feature vectors of each historical video to determine the information expressed by each historical video; and then organize the obtained feature information to form an input set of the neural network model; A model building unit, used to build a neural network model, train and test the neural network model using an input set, and obtain a model that can automatically recognize video content; The model practice unit is used to apply the tested neural network model to the current video to automatically identify the feature content of the current video and the information expressed by the video content.
9. The artificial intelligence-based video content processing system according to claim 8, characterized in that: Also includes: A human-computer interaction terminal is used to receive video content uploaded by users and visualize the processing results of the video content; Suggestion feedback unit: users use the human-computer interaction terminal as a carrier to provide the system with suggestions and opinions on the video content processing results; if the user has objections to the processing results of the current video content, the system will re-identify and process the current video content.
10. The artificial intelligence-based video content processing system according to claim 8, characterized in that: Also includes: The progress prompt unit is used to send a progress prompt message to the user after the system outputs the processing result of the video content.
Citation Information
Patent Citations
Video frequency object tracking method, device and automatic video frequency following system
CN101290681B
Video labeling method
CN101827203A
Automatic video advertisement detection method
CN103605991A
Voice noise-reducing method and voice noise-reducing device
CN103700375A
Video content batch supervision cluster system and method
CN111246168A