Method and system for detecting materials before casting

By extracting the color and texture characteristics of video materials and combining with neural network models, the problem of inefficient traditional manual detection is solved, intelligent material compliance and audience matching evaluation is achieved, and the delivery effect of creatives is improved.

CN120302085APending Publication Date: 2025-07-11JIANGXI AOXING BRILLIANT NETWORK TECH CO LTD

Patent Information

Application Number
CN202510380510.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Traditional manual detection of creatives is inefficient, prone to missed reviews and misjudgments, and the inability to effectively evaluate the compatibility between the video material and the target audience, resulting in poor creative delivery.

Method used

By obtaining video frames of video materials, extracting color and texture features, analyzing scene conversion information, performing image recognition and speech recognition, and combining neural network models to output personalized delivery plans, improving material compliance and matching.

Benefits of technology

It realizes the compliance and audience matching of intelligent video materials, and improves the delivery effect and efficiency of creatives.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120302085A_ABST
    Figure CN120302085A_ABST
Patent Text Reader

Abstract

The invention provides a material pre-projection detection method and system, and the method comprises the steps: obtaining a video material, carrying out the video frame extraction of the video material, and determining the color features and texture features of each frame of image; analyzing scene conversion information in the video material according to the color features and the texture features, determining an image of a corresponding scene, and performing image recognition; judging whether the corresponding image is compliant or not; if yes, determining product category information of the video material; extracting audios in the video materials, determining volume abnormal information according to the audios, performing voice recognition on the audios, and converting the audios into texts; judging whether the text is compliant or not; if yes, calculating the similarity of the adjacent frames of images, and determining lens switching frequency information according to the similarity; and inputting the delivery area information, the product category information, the volume abnormal information, the lens switching frequency information and the scene conversion information into the trained neural network model, outputting a delivery plan, and finally achieving the purpose of improving the material delivery effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of pre - investment detection of materials, and particularly relates to a method and system for pre - investment detection of materials. Background Art

[0002] Advertising materials need to be detected before being put on the market. The detection mainly focuses on the compliance, quality, and the matching degree of style and positioning of the materials. Traditional manual detection relies on the subjective judgment and attention concentration of the detection personnel, which is prone to situations such as missed examination and misjudgment, and has low efficiency.

[0003] Especially for video materials, it is necessary to detect three dimensions: images, audio, and time series. For manual detection, it is time - consuming and laborious. In addition, there is currently no good way to evaluate the fit between a video material and the target audience, resulting in a great reduction in the advertising material placement effect. Summary of the Invention

[0004] Based on this, embodiments of the present invention provide a method and system for pre - investment detection of materials, aiming to perform intelligent detection on video materials, and make a placement plan for the target area according to the detection results, so as to improve the material placement effect.

[0005] The first aspect of the embodiments of the present invention provides a method for pre - investment detection of materials, which is applied to the scenario of elevator video advertisement placement. The method includes:

[0006] Obtain a video material, extract video frames from the video material, and determine the color features and texture features of each frame of the extracted image;

[0007] Analyze the scene transition information in the video material according to the color features and the texture features;

[0008] Determine the image corresponding to the scene according to the scene transition information, and perform image recognition;

[0009] Judge whether the corresponding image is compliant according to the image recognition result;

[0010] If the image is compliant, determine the product category information of the video material according to the image recognition result;

[0011] If the image is non - compliant, terminate the material detection and give an early warning;

[0012] Extract the audio in the video material, determine the abnormal volume information according to the audio, and perform speech recognition on the audio to convert it into text;

[0013] Judge whether the text is compliant;

[0014] If the text is compliant, calculate the similarity between adjacent frame images, and determine the shot transition frequency information according to the similarity;

[0015] If the text is non-compliant, terminate the material detection and give an alarm;

[0016] Input the placement area information, the product category information, the volume anomaly information, the shot transition frequency information, and the scene transition information into the trained neural network model to output the placement plan.

[0017] Further, the step of determining the color features and texture features of each frame image according to the extracted frame images includes:

[0018] Separate the three RGB color channels of each frame image, and calculate the histogram of each color channel according to the separated color channels;

[0019] Merge the histograms of the three color channels to obtain a merged result, and the merged result is the color feature;

[0020] Convert each frame image into a grayscale image, and calculate the gray-level co-occurrence matrix of the grayscale image;

[0021] Determine the texture feature according to the gray-level co-occurrence matrix.

[0022] Further, the step of analyzing the scene transition information in the video material according to the color feature and the texture feature includes:

[0023] Adopt the principal component analysis method to perform dimensionality reduction processing on the color feature and the texture feature respectively;

[0024] Perform clustering analysis on the dimensionality-reduced color feature and texture feature respectively, and determine the scene transition information according to the change of the clustering result over time, where the scene transition information at least includes the scene category, the time point of scene transition, and the scene categories before and after the transition;

[0025] Compare the scene transition information analyzed according to the color feature and the texture feature, determine the different parts, and check the different parts to determine the final scene transition information.

[0026] Further, the step of extracting the audio in the video material and determining the volume anomaly information according to the audio includes:

[0027] Obtain the audio signal, and perform Fourier transform on the audio signal to convert the audio signal from the time domain to the frequency domain;

[0028] Based on the results of the Fourier transform, determine the energy distribution of different frequency components in the audio signal and calculate the root mean square energy;

[0029] Determine whether the root mean square energy is greater than the first threshold;

[0030] If it is determined that the root mean square energy is greater than the first threshold, it is determined that there is an abnormal volume fluctuation, and obtain the number of abnormal volume fluctuations and the corresponding abnormal volume fluctuation time.

[0031] Further, the step of calculating the similarity of adjacent frame images and determining the shot transition frequency information according to the similarity includes:

[0032] Obtain the pixel points and corresponding pixel values of adjacent frame images respectively, and calculate the MSE value by the mean square error method;

[0033] Determine whether the MSE value is greater than the first threshold;

[0034] If it is determined that the MSE value is greater than the first threshold, it is determined that a shot transition has occurred, and count the number of shot transitions per unit time.

[0035] Further, the neural network model is composed of an input layer, a hidden layer and an output layer. The input layer contains a number of input nodes, and the number of input nodes is the same as the number of types of input data, which is used to receive different types of input data;

[0036] The hidden layer consists of five sub-layers. The first hidden sub-layer has 64 nodes, and uses the LeakyReLU function to extract features from the input data to find the key features related to the output result. The second and third hidden sub-layers each have 128 nodes, and use the Sigmoid function to further mine the features of the output data of the first hidden sub-layer and capture the internal connections between nodes. The fourth and fifth hidden sub-layers each contain 64 nodes, and with the help of the ReLU activation function, perform secondary processing on the data features output by the third hidden sub-layer;

[0037] The output layer is provided with a number of output nodes, and the number of output nodes corresponds to the number of types of output data. The output layer uses a linear activation function to output the final prediction result.

[0038] Further, in the step of inputting the placement area information, the product category information, the volume anomaly information, the shot transition frequency information and the scene conversion information into the trained neural network model and outputting the placement plan, the placement area information at least includes the number of people in different time periods of the placement area, the user characteristics of the placement area and the user preferences of the placement area, and the placement plan at least includes the target elevator for placement and the placement time period.

[0039] The second aspect of the embodiments of the present invention provides a pre-delivery detection system for a material, which is used to implement the pre-delivery detection method for a material provided in the first aspect of the embodiments of the present invention. The system includes:

[0040] A first determination module, configured to obtain a video material, extract video frames from the video material, and determine the color features and texture features of each frame image according to the extracted frame images;

[0041] An analysis module, configured to analyze the scene transition information in the video material according to the color features and the texture features;

[0042] A second determination module, configured to determine the image corresponding to the scene according to the scene transition information and perform image recognition;

[0043] A first judgment module, configured to judge whether the corresponding image is compliant according to the image recognition result;

[0044] A third determination module, configured to, if the image is compliant, determine the product category information of the video material according to the image recognition result;

[0045] A first warning module, configured to, if the image is non-compliant, terminate the material detection and give a warning;

[0046] A fourth determination module, configured to extract the audio in the video material, determine the volume abnormality information according to the audio, and perform speech recognition on the audio to convert it into text;

[0047] A second judgment module, configured to judge whether the text is compliant;

[0048] A fifth determination module, configured to, if the text is compliant, calculate the similarity between adjacent frame images, and determine the shot transition frequency information according to the similarity;

[0049] A second warning module, configured to, if the text is non-compliant, terminate the material detection and give a warning;

[0050] An input module, configured to input the placement area information, the product category information, the volume abnormality information, the shot transition frequency information, and the scene transition information into a trained neural network model, and output a placement plan.

[0051] The third aspect of the embodiments of the present invention provides a computer-readable storage medium, including:

[0052] The readable storage medium stores one or more programs, which when executed by a processor, implement the pre-delivery detection method for a material as described in the first aspect.

[0053] The fourth aspect of the embodiments of the present invention provides an electronic device, which includes a memory and a processor, wherein:

[0054] The memory is used to store computer programs;

[0055] When the processor is used to execute the computer program stored in the memory, the pre-delivery detection method for materials as described in the first aspect is implemented.

[0056] A pre-delivery detection method and system for materials provided in an embodiment of the present invention, by acquiring video materials, extracting video frames from the video materials, determining the color features and texture features of each frame of the extracted images according to the extracted frame images; analyzing the scene transition information in the video materials according to the color features and texture features; determining the images corresponding to the scenes according to the scene transition information, and performing image recognition; judging whether the corresponding images are compliant according to the image recognition results; if so, determining the product category information of the video materials according to the image recognition results; extracting the audio in the video materials, determining the volume anomaly information according to the audio, and performing speech recognition on the audio to convert it into text; judging whether the text is compliant; if so, calculating the similarity of adjacent frame images, and determining the shot transition frequency information according to the similarity; inputting the placement area information, product category information, volume anomaly information, shot transition frequency information, and scene transition information into a trained neural network model, and outputting a placement plan. Specifically, through the above method, while intelligently auditing video materials, a personalized placement plan can be given to achieve the purpose of improving the placement effect of materials. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 It is a flowchart of the implementation of a pre-delivery detection method for materials provided in Embodiment 1 of the present invention;

[0058] Figure 2 It is a structural block diagram of a pre-delivery detection method system provided in Embodiment 2 of the present invention;

[0059] Figure 3 It is a structural block diagram of an electronic device provided in Embodiment 3 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0060] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.

[0061] It should be noted that when an element is referred to as being "fixed to" another element, it can be directly on the other element or there can also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are for illustrative purposes only.

[0062] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this invention belongs. The terms used herein in the specification of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the associated listed items.

[0063] Embodiment 1

[0064] Embodiment 1 of the present invention provides a method for pre - detection of materials, which is applied to the scenario of elevator video advertisement placement. Please refer to Figure 1 , which is a flowchart for the implementation of a method for pre - detection of materials, specifically including steps S01 to S11.

[0065] Step S01: Obtain a video material, extract video frames from the video material, and determine the color features and texture features of each frame image according to the extracted frame images.

[0066] Specifically, first determine the original frame rate of the video material. For example, common video frame rates are 24 frames per second, 30 frames per second, 60 frames per second, etc. According to actual requirements and computing resources, determine a fixed extraction interval. Taking a video with 30 frames per second as an example, extract one frame every 10 frames, that is, extract 3 frames per second. In this way, it can retain the main information of the video to a certain extent and reduce the computational amount of subsequent processing. In the embodiment of the present invention, the VideoCapture function of the OpenCV library can be used to open the video file, and the frame image to be extracted can be located by setting the frame number position (CAP_PROP_POS_FRAMES) of the video, and then the read function is used to read the frame image.

[0067] To further reduce the amount of image analysis on this basis, after extracting each frame of the image and determining the color features and texture features of each frame of the image, it should be noted that the RGB three color channels of each frame of the image are separated, and according to the separated color channels, the histogram of each color channel is calculated. Among them, for each extracted frame of the image, the split function of OpenCV is used to separate the RGB three color channels of the image. This is because the pixel distribution of different color channels can reflect the color features of the image. For example, the distribution of pixel values in the red channel can reflect the intensity and distribution of the red tone in the image. And for each separated color channel, the calcHist function of OpenCV is used to calculate the color histogram. This function will count the number of occurrences of different pixel values in each color channel, so as to obtain the pixel distribution of this channel. The color space can be quantized into 256 levels, so that the histogram of each color channel has 256 bins, and each bin represents a pixel value range;

[0068] The histograms of the three color channels are combined to obtain a combined result, and the combined result is the color feature. It can be understood that after calculating the histograms of each color channel (R, G, B), since each histogram is a one-dimensional vector, they can be directly concatenated in order to form a longer one-dimensional vector. Among them, before concatenation, the histogram of each channel is normalized. This can eliminate the difference in the number of pixels between different channels and make the feature vector more comparable. The normalization method can be to divide the histogram elements of each channel by the total number of pixels in this channel, or use L1 or L2 normalization. In addition, when concatenating, a weighted concatenation method can be adopted, that is, according to the importance of different channels in image analysis, different weights are assigned to the histograms of each channel, and then concatenation is performed;

[0069] Each frame of the image is converted into a grayscale image, and the gray-level co-occurrence matrix of the grayscale image is calculated. Specifically, the cvtColor function of OpenCV can be used to convert the RGB image into a grayscale image, which can simplify the subsequent texture analysis process. Then the calcGLCM function of OpenCV is used to calculate the gray-level co-occurrence matrix. Among them, the gray-level co-occurrence matrix describes the gray value distribution relationship of pixel pairs in the image at a certain distance and angle. By analyzing the gray-level co-occurrence matrix at different distances and angles, the texture details of the image can be obtained;

[0070] According to the gray-level co-occurrence matrix, the texture feature is determined. The texture feature can be contrast, correlation, energy, homogeneity, etc.

[0071] Step S02, analyze the scene transition information in the video material according to the color feature and the texture feature.

[0072] It should be noted that the principal component analysis (PCA) is used to reduce the dimensionality of the color feature and the texture feature respectively to reduce the amount of analysis;

[0073] The color feature and the texture feature after dimensionality reduction are respectively subjected to K-Means clustering analysis, and the scene transition information is determined according to the change of the clustering result over time. It can be understood that the K-Means algorithm will cluster similar video frames into one category, and different clusters represent different scenes. Among them, the scene transition information at least includes the scene category, the time point of scene transition, and the scene categories before and after the transition. It can be understood that by observing the change of the clustering result over time, the scene transition situation is determined. When a video frame is converted from one clustering category to another, it is determined that a scene transition has occurred, and the time point of scene transition and the scene categories before and after the transition are recorded for subsequent analysis;

[0074] The scene transition information obtained by analyzing the color feature and the texture feature is compared to determine the different parts, and the different parts are verified to determine the final scene transition information. Specifically, it includes the comparison of the number of scene categories, the time point of scene transition, and the scene categories before and after the transition. First, compare whether the number of scene categories is the same. If the number of scene categories obtained by analyzing the color feature and the texture feature is both 3, it means that the number of scene categories is the same. If the comparison of the number of scene categories is the same, then compare whether the scene categories before and after the transition are the same. Suppose the scene categories obtained by analyzing the color feature and the texture feature both include Scene A, Scene B, and Scene C. The order of the scene categories obtained by analyzing the color feature is Scene A, Scene B, and Scene C, while the order of the scene categories obtained by analyzing the texture feature is Scene A, Scene C, and Scene B. Then it means that the scene categories before and after the transition are different. If the comparison of the scene categories before and after the transition is the same, then compare whether the time points of scene transition are the same. Among them, when the difference between the time points of scene transition is within a certain range, it is regarded as the same. If there is no difference in the comparison result, the scene transition information is input into the subsequent processing steps. If there is a difference in the comparison result, it can be identified manually, and the analysis timing of the color feature and the texture feature in the video material can be determined by machine learning, which provides help for accurately obtaining the scene transition information of the video material in the future.

[0075] Step S03, according to the scene transition information, determine the image corresponding to the scene and perform image recognition.

[0076] In this embodiment, since the scene conversion information is obtained, the corresponding images can be extracted according to the scene category and the images can be recognized, which can reduce the analysis amount again. Specifically, a suitable object detection algorithm based on deep learning can be selected to obtain the object information in the images, and the object detection algorithm can use the YOLO series (YOLOv5, YOLOv8, etc.).

[0077] Step S04: According to the image recognition result, determine whether the corresponding image is compliant. If so, execute Step S05; if not, execute Step S06.

[0078] Among them, according to the obtained object information, compare it with the relevant regulations to determine whether it is compliant.

[0079] Step S05: Then, according to the image recognition result, determine the product category information of the video material.

[0080] It can be understood that the product categories include electronic products, daily necessities, food, clothing, etc.

[0081] Step S06: Then terminate the material detection and give an alarm.

[0082] Step S07: Extract the audio in the video material, determine the volume anomaly information according to the audio, and perform speech recognition on the audio to convert it into text.

[0083] Specifically, first obtain the audio signal, and perform Fourier transform on the audio signal to convert the audio signal from the time domain to the frequency domain. It should be noted that the load function of the Librosa library is used to load the audio part in the video, and then the stft (short-time Fourier transform) function of Librosa is used to convert the audio signal from the time domain to the frequency domain. The short-time Fourier transform can divide the audio signal into multiple short time periods and perform Fourier transform on each time period, so as to obtain the change of the energy distribution of the audio signal at different frequency components over time;

[0084] According to the result of the Fourier transform, determine the energy distribution of different frequency components in the audio signal and calculate the root mean square energy. Among them, the feature.rms function of Librosa is used to calculate the root mean square energy of the audio signal;

[0085] Judge whether the root mean square energy is greater than the first threshold;

[0086] If it is judged that the root mean square energy is greater than the first threshold, it is determined that there is an abnormal volume fluctuation, and the number of abnormal volume fluctuations and the corresponding abnormal volume fluctuation time are obtained.

[0087] Further, through speech recognition technology (such as Baidu Speech Recognition API), the speech in the audio can be converted into text. Before inputting the audio into the speech recognition API, some preprocessing can be performed on the audio, such as removing noise and adjusting the sampling rate, etc., to improve the recognition accuracy.

[0088] Step S08, determine whether the text is compliant. If it is, then execute Step S09. If not, then execute Step S10.

[0089] Step S09, if the text is compliant, then calculate the similarity of adjacent frame images, and determine the shot transition frequency information according to the similarity.

[0090] Specifically, obtain the pixel points and corresponding pixel values of adjacent frame images respectively, and calculate the MSE value through the mean squared error method. The mean squared error refers to the average of the squares of the differences between the corresponding pixel values of two frames of images. The smaller the MSE value, the more similar the two frames of images are. The meanSquaredError function of OpenCV can be used to calculate the MSE value;

[0091] Judge whether the MSE value is greater than the first threshold;

[0092] If it is judged that the MSE value is greater than the first threshold, then it is determined that a shot transition has occurred, and the number of shot transitions within a unit time is counted. It can be understood that by analyzing the shot transition frequency, the rhythm of the video can be understood. For example, a fast-paced video usually has a higher shot transition frequency.

[0093] Step S10, then terminate the material detection and give an alarm.

[0094] Step S11, input the placement area information, the product category information, the volume anomaly information, the shot transition frequency information, and the scene transition information into the trained neural network model, and output the placement plan.

[0095] In this embodiment, the neural network model is composed of an input layer, a hidden layer, and an output layer. The input layer contains a number of input nodes, and the number of input nodes is the same as the number of types of input data, which is used to receive different types of input data;

[0096] The hidden layer is composed of five sub-layers. The first hidden sub-layer has 64 nodes, and the LeakyReLU function is used to extract the features of the input data to find the key features related to the output result. The second hidden sub-layer and the third hidden sub-layer each have 128 nodes, and the Sigmoid function is used to further explore the features of the output data of the first hidden sub-layer and capture the internal connections between the nodes. The fourth hidden sub-layer and the fifth hidden sub-layer each contain 64 nodes, and with the help of the ReLU activation function, the data features output by the third hidden sub-layer are processed twice;

[0097] The output layer is provided with a number of output nodes, and the number of output nodes corresponds to the number of output data types. The output layer uses a linear activation function to output the final prediction result.

[0098] Specifically, 80% of the dataset samples after random shuffling are used as training set samples, and the other 20% are used as test set samples. Among them, the neuron formulas of the hidden layer and the output layer can be expressed as:

[0099] The first hidden sub-layer:

[0100]

[0101] The second hidden sub-layer:

[0102]

[0103] The third hidden sub-layer:

[0104]

[0105] The fourth hidden sub-layer:

[0106]

[0107] The fifth hidden sub-layer:

[0108]

[0109] The output layer:

[0110]

[0111] Among them, x j and y i respectively represent the j-th node of the input layer and the i-th node of the output layer, ω l,i,j represents the weight connecting the j-th node of the (l - 1)-th layer and the i-th node of the output layer, b l,i represents the bias term of the i-th node of the l-th layer. Exemplarily, h 1,i represents the i-th neuron of the first hidden sub-layer, h 2,i represents the i-th neuron of the second hidden sub-layer, h 3,i represents the i-th neuron of the third hidden sub-layer, h 4,i represents the i-th neuron of the fourth hidden sub-layer, h 5,i represents the i-th neuron of the fifth hidden sub-layer. Similarly, h 1,j represents the j-th neuron of the first hidden sub-layer, h 2,j represents the j-th neuron of the second hidden sub-layer, h 3,j represents the j-th neuron of the third hidden sub-layer, h4,j The j-th neuron represented as the fourth hidden sub-layer, h 5,j The j-th neuron represented as the fifth hidden sub-layer. LeakyReLU, Sigmoid, and ReLU are activation functions.

[0112] It can be understood that during model training, the input data at least includes the number of people flow at different time periods in the placement area, user characteristics in the placement area, user preferences in the placement area, product categories, scene categories, time points of scene conversion, scene categories before and after conversion, the number of abnormal volume fluctuations, corresponding abnormal volume fluctuation times, and the number of camera cuts per unit time. The output data at least includes the target elevator for placement, placement time period, and economic benefits generated. Among them, the above input data and output data are all historical data, and the data corresponding to successful cases can be input into the model for training.

[0113] It should be noted that the economic benefits generated are mainly counted as the economic benefits generated in the placement area from the start of placement to the end of placement. In addition, the number of people flow at different time periods in the placement area, user characteristics in the placement area, and user preferences in the placement area can be obtained through big data. In addition, user characteristics can be understood as that the audience of community elevators may mainly be residents, and the audience of office building elevators are mostly office workers. User preferences can be understood as that community residents may be more concerned about life and family products, and the advertising content and style can be close to life scenes. For areas with a large number of young audiences, the advertising style can be more fashionable, lively, and incorporate popular elements.

[0114] By inputting relevant parameters into the trained neural network model, a placement plan can be output to improve the placement effect of the material.

[0115] In summary, a method for pre-detection of materials proposed in an embodiment of the present invention. This method obtains video materials, extracts video frames from the video materials, determines the color features and texture features of each frame image according to the extracted frame images; analyzes the scene conversion information in the video materials according to the color features and texture features; determines the images corresponding to the scenes according to the scene conversion information, and performs image recognition; determines whether the corresponding images are compliant according to the image recognition results; if so, determines the product category information of the video materials according to the image recognition results; extracts the audio in the video materials, determines the abnormal volume information according to the audio, and performs speech recognition on the audio to convert it into text; determines whether the text is compliant; if so, calculates the similarity of adjacent frame images, and determines the camera cut frequency information according to the similarity; inputs the placement area information, product category information, abnormal volume information, camera cut frequency information, and scene conversion information into the trained neural network model to output a placement plan. Specifically, through the above method, while intelligently reviewing video materials, a personalized placement plan can be given to achieve the purpose of improving the placement effect of the materials.

[0116] Example 2

[0117] Example 2 of the present invention provides a pre-material detection system 200. Please refer to Figure 2 , which is a structural block diagram of a pre-material detection method system. The pre-material detection system 200 includes:

[0118] A first determination module 201, configured to obtain a video material, extract video frames from the video material, and determine the color feature and texture feature of each frame image according to the extracted frame images;

[0119] An analysis module 202, configured to analyze the scene transition information in the video material according to the color feature and the texture feature;

[0120] A second determination module 203, configured to determine the image corresponding to the scene according to the scene transition information and perform image recognition;

[0121] A first judgment module 204, configured to judge whether the corresponding image is compliant according to the image recognition result;

[0122] A third determination module 205, configured to, if the image is compliant, determine the product category information of the video material according to the image recognition result;

[0123] A first warning module 206, configured to, if the image is non-compliant, terminate the material detection and give a warning;

[0124] A fourth determination module 207, configured to extract the audio in the video material, determine the volume anomaly information according to the audio, and perform speech recognition on the audio to convert it into text;

[0125] A second judgment module 208, configured to judge whether the text is compliant;

[0126] A fifth determination module 209, configured to, if the text is compliant, calculate the similarity of adjacent frame images, and determine the shot transition frequency information according to the similarity;

[0127] A second warning module 210, configured to, if the text is non-compliant, terminate the material detection and give a warning;

[0128] An input module 211, configured to input the placement area information, the product category information, the volume anomaly information, the shot transition frequency information, and the scene transition information into a trained neural network model, and output a placement plan. The neural network model is composed of an input layer, a hidden layer, and an output layer. The input layer includes a plurality of input nodes, and the number of input nodes is the same as the number of types of input data, and is used to receive different types of input data;

[0129] The hidden layer consists of five sub-layers. The first hidden sub-layer has 64 nodes, which extracts features from the input data through the LeakyReLU function to find the key features related to the output result. The second and third hidden sub-layers each have 128 nodes, and the Sigmoid function is used to further explore the features of the output data of the first hidden sub-layer and capture the internal connections between nodes. The fourth and fifth hidden sub-layers each contain 64 nodes, and with the help of the ReLU activation function, the data features output by the third hidden sub-layer are processed twice;

[0130] The output layer is provided with a number of output nodes, and the number of output nodes corresponds to the number of types of output data. The output layer uses a linear activation function to output the final prediction result;

[0131] The information of the placement area at least includes the number of people flow in different time periods of the placement area, the user characteristics of the placement area, and the user preferences of the placement area. The placement plan at least includes the target elevators for placement and the placement time period.

[0132] Further, in some other embodiments of the present invention, the first determination module 201 includes:

[0133] A first calculation unit, which is used to separate the RGB three color channels of each frame of image, and calculate the histogram of each color channel according to the separated color channels;

[0134] A merging unit, which is used to merge the histograms of the three color channels to obtain a merging result, and the merging result is the color feature;

[0135] A second calculation unit, which is used to convert each frame of image into a grayscale image and calculate the gray-level co-occurrence matrix of the grayscale image;

[0136] A first determination unit, which is used to determine the texture feature according to the gray-level co-occurrence matrix.

[0137] Further, in some other embodiments of the present invention, the analysis module 202 includes:

[0138] A dimensionality reduction processing unit, which is used to perform dimensionality reduction processing on the color feature and the texture feature respectively by using the principal component analysis method;

[0139] A clustering analysis unit, which is used to perform clustering analysis on the dimensionality-reduced color feature and texture feature respectively, and determine the scene conversion information according to the change of the clustering result over time. Among them, the scene conversion information at least includes the scene category, the time point of scene conversion, and the scene categories before and after conversion;

[0140] A comparison unit is configured to compare the scene conversion information obtained by analyzing the color feature and the texture feature, determine the different parts, and check the different parts to determine the final scene conversion information.

[0141] Further, in some other embodiments of the present invention, the fourth determination module 207 includes:

[0142] A Fourier transform unit is configured to obtain an audio signal and perform a Fourier transform on the audio signal to convert the audio signal from the time domain to the frequency domain;

[0143] A second determination unit is configured to determine the energy distribution of different frequency components in the audio signal according to the result of the Fourier transform and calculate the root mean square energy;

[0144] A first judgment unit is configured to judge whether the root mean square energy is greater than a first threshold;

[0145] A third determination unit is configured to, if it is judged that the root mean square energy is greater than the first threshold, determine that there is an abnormal fluctuation in volume and obtain the number of abnormal volume fluctuations and the corresponding abnormal volume fluctuation time.

[0146] Further, in some other embodiments of the present invention, the fifth determination module 209 includes:

[0147] A third calculation unit is configured to respectively obtain the pixel points and the corresponding pixel values of adjacent frame images and calculate the MSE value by means of the mean square error method;

[0148] A second judgment unit is configured to judge whether the MSE value is greater than a first threshold;

[0149] A fourth determination unit is configured to, if it is judged that the MSE value is greater than the first threshold, determine that a shot change has occurred and count the number of shot changes per unit time.

[0150] Embodiment III

[0151] Embodiment III of the present invention proposes an electronic device. Please refer to Figure 3 , which is a structural block diagram of an electronic device, including a memory 20, a processor 10, and a computer program 30 stored in the memory and executable on the processor. When the processor 10 executes the computer program 30, the method for pre-detection of materials as described above is implemented.

[0152] Among them, in some embodiments, the processor 10 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips, and is configured to run the program code stored in the memory 20 or process data, such as executing an access restriction program, etc.

[0153] Among them, the memory 20 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. The memory 20 can be an internal storage unit of the electronic device in some embodiments, such as the hard disk of the electronic device. The memory 20 can also be an external storage device of the electronic device in other embodiments, such as a plug-in hard disk equipped on the electronic device, a Smart Media Card (SMC), a Secure Digital (SD) card, a FlashCard, etc. Further, the memory 20 can also include both an internal storage unit and an external storage device of the electronic device. The memory 20 can be used not only to store application software and various types of data of the electronic device, but also to temporarily store data that has been output or will be output.

[0154] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the material pre-detection method as described above.

[0155] Those skilled in the art can understand that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus or device and execute the instructions), or in combination with these instruction execution systems, apparatus or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate or transport a program for use by or in connection with an instruction execution system, apparatus or device.

[0156] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part (electronic device) having one or more wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting or otherwise processing it as appropriate, and then storing it in a computer memory.

[0157] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one of the following techniques known in the art or a combination thereof can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0158] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0159] The above embodiments merely represent several implementation manners of the present invention, and the description thereof is relatively specific and detailed, but should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the appended claims.

Claims

1. A method for pre - investment detection of materials, characterized in that, Applied to the scenario of elevator video advertisement placement, the method includes: Obtain video materials, extract video frames from the video materials, and determine the color features and texture features of each frame image according to the extracted frame images; Analyze the scene transition information in the video materials according to the color features and the texture features; Determine the images corresponding to the scenes according to the scene transition information and perform image recognition; Judge whether the corresponding images are compliant according to the image recognition results; If the images are compliant, determine the product category information of the video materials according to the image recognition results; If the images are not compliant, terminate the material detection and give an alarm; Extract the audio in the video materials, determine the volume anomaly information according to the audio, and perform speech recognition on the audio to convert it into text; Judge whether the text is compliant; If the text is compliant, calculate the similarity of adjacent frame images, and determine the shot transition frequency information according to the similarity; If the text is not compliant, terminate the material detection and give an alarm; Input the placement area information, the product category information, the volume anomaly information, the shot transition frequency information, and the scene transition information into the trained neural network model to output a placement plan.

2. The pre-investment detection method for materials according to claim 1, wherein The step of determining the color features and texture features of each frame image according to the extracted frame images includes: Separate the three RGB color channels of each frame image, and calculate the histogram of each color channel according to the separated color channels; Merge the histograms of the three color channels to obtain a merged result, and the merged result is the color feature; Convert each frame image into a grayscale image and calculate the gray-level co-occurrence matrix of the grayscale image; Determine the texture feature according to the gray-level co-occurrence matrix.

3. The pre-investment detection method for materials according to claim 2, characterized in that The step of analyzing the scene transition information in the video materials according to the color features and the texture features includes: Adopt the principal component analysis method to perform dimensionality reduction processing on the color features and the texture features respectively; Perform clustering analysis on the dimensionality-reduced color features and texture features respectively, and determine the scene transition information according to the change of the clustering results over time, where the scene transition information at least includes the scene category, the time point of scene transition, and the scene categories before and after the transition; Compare the scene transition information analyzed according to the color features and the texture features, determine the different parts, and check the different parts to determine the final scene transition information.

4. The material pre-investment detection method according to claim 3, characterized in that The step of extracting the audio in the video materials and determining the volume anomaly information according to the audio includes: Obtain an audio signal, and perform Fourier transform on the audio signal to convert the audio signal from the time domain to the frequency domain; Determine the energy distribution of different frequency components in the audio signal according to the result of Fourier transform, and calculate the root mean square energy; Judge whether the root mean square energy is greater than a first threshold; If it is judged that the root mean square energy is greater than the first threshold, determine that there is abnormal volume fluctuation, and obtain the number of abnormal volume fluctuations and the corresponding abnormal volume fluctuation time.

5. The material pre-investment detection method according to claim 4, characterized in that The step of calculating the similarity of adjacent frame images and determining the shot transition frequency information according to the similarity includes: Obtain the pixel points and corresponding pixel values of adjacent frame images respectively, and calculate the MSE value by the mean square error method; Judge whether the MSE value is greater than the first threshold; If it is judged that the MSE value is greater than the first threshold, it is determined that a shot transition has occurred, and the number of shot transitions within a unit time is counted.

6. The pre-investment detection method for materials according to claim 5, wherein The neural network model consists of an input layer, a hidden layer and an output layer. The input layer contains a number of input nodes, and the number of input nodes is the same as the number of types of input data, which is used to receive different types of input data; The hidden layer consists of five sub-layers. The first hidden sub-layer has 64 nodes, and the LeakyReLU function is used to extract features from the input data to find the key features related to the output result. The second hidden sub-layer and the third hidden sub-layer each have 128 nodes, and the Sigmoid function is used to further explore the features of the output data of the first hidden sub-layer and capture the internal connections between nodes. The fourth hidden sub-layer and the fifth hidden sub-layer each contain 64 nodes, and with the help of the ReLU activation function, the data features output by the third hidden sub-layer are processed twice; The output layer is provided with a number of output nodes, and the number of output nodes corresponds to the number of types of output data. The output layer uses a linear activation function to output the final prediction result.

7. The pre-investment detection method for materials according to claim 6, wherein In the step of inputting the placement area information, the product category information, the volume anomaly information, the shot transition frequency information and the scene conversion information into the trained neural network model to output the placement plan, the placement area information at least includes the number of people in the placement area at different time periods, the user characteristics of the placement area and the user preferences of the placement area, and the placement plan at least includes the target elevator to be placed and the placement time period.

8. A pre-investment detection system for materials, characterized in that, For implementing the pre-detection method of materials as described in any one of claims 1-7, the system includes: A first determination module, configured to obtain a video material, extract video frames from the video material, and determine the color features and texture features of each frame image according to the extracted frame images; An analysis module, configured to analyze the scene conversion information in the video material according to the color features and the texture features; A second determination module, configured to determine the image corresponding to the scene according to the scene conversion information and perform image recognition; A first judgment module, configured to judge whether the corresponding image is compliant according to the image recognition result; A third determination module, configured to, if the image is compliant, determine the product category information of the video material according to the image recognition result; A first warning module, configured to, if the image is non-compliant, terminate the material detection and give a warning; A fourth determination module, configured to extract the audio in the video material, determine the volume anomaly information according to the audio, and perform speech recognition on the audio to convert it into text; A second judgment module, configured to judge whether the text is compliant; A fifth determination module, configured to, if the text is compliant, calculate the similarity of adjacent frame images, and determine the shot transition frequency information according to the similarity; A second warning module, configured to, if the text is non-compliant, terminate the material detection and give a warning; An input module, configured to input the placement area information, the product category information, the volume anomaly information, the lens switching frequency information, and the scene conversion information into a trained neural network model, and output a placement plan.

9. A computer-readable storage medium, characterized in that, Comprising: The readable storage medium stores one or more programs, which when executed by a processor implement the pre-placement detection method of the material according to any one of claims 1-7.

10. An electronic device, characterized in that, The electronic device includes a memory and a processor, wherein: The memory is used for storing a computer program; When the processor executes the computer program stored on the memory, it implements the pre-placement detection method of the material according to any one of claims 1-7.

Citation Information

Patent Citations

  • An artificial intelligence display screen advertisement dynamic delivery system and method

    CN109934625A

  • Video scene recognition method and device, storage medium and electronic device

    CN110147711A

  • Video fine structuring method based on multi-feature fusion

    CN110188625A

  • Advertisement sensitive content auditing method and system based on artificial intelligence

    CN116415017A

  • Image recognition method and device, equipment and storage medium

    CN116977684A

Cited By

  • Advertisement risk analysis method and device, equipment and storage medium

    CN121305421A