A method for determining low-quality video materials of an operator based on matching degrees of element objects
By calculating the matching degree between audio text elements and visual elements in video clips, as well as the matching degree between thematic keywords of the material, the problem of inaccurate evaluation of low-quality video material in existing technologies is solved, and a more comprehensive evaluation of video material quality is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- E-JOINED INTERNET & TECH CO LTD
- Filing Date
- 2026-02-06
- Publication Date
- 2026-04-28
AI Technical Summary
In existing technologies, operators' results in identifying low-quality video footage are incomplete and inaccurate, relying on preset video footage templates, which leads to inaccurate evaluations.
By calculating the matching degree between audio text elements and visual elements in video clips, and combining it with the matching degree of thematic keywords, segment and overall analysis is performed to evaluate the quality of video materials, including video segmentation, semantic recognition of audio text elements and visual elements, vector construction, and keyword matching.
It improves the comprehensiveness and accuracy of the results in identifying low-quality video footage, and enables video footage quality assessment based on segment and overall analysis.
Smart Images

Figure CN121661575B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of data processing technology, specifically relating to a method for identifying low-quality video materials from operators based on element object matching degree. Background Technology
[0002] With the development of the internet and multimedia communications, using multimedia information such as images, audio, and video for product demonstrations has become a mainstream promotional method in various fields, especially for telecom operators using video materials to promote newly launched products. To ensure that operators' product information is accurately displayed, data processing of produced video materials to analyze their quality has become a pressing issue in the industry.
[0003] In existing technologies, video material templates corresponding to various product types are pre-set. The video material to be evaluated is then compared with the video style and necessary video elements in the corresponding templates. Video material whose comparison results exceed a preset fluctuation range is identified as low-quality video material. However, existing technologies rely on preset video material templates, resulting in incomplete and inaccurate identification of low-quality video material. Summary of the Invention
[0004] This application provides a method for identifying low-quality video materials from operators based on element object matching degree. It solves the problems of incomplete and inaccurate identification results of low-quality video materials in the prior art. By calculating the first matching degree between audio text elements and visual elements in each video segment, the segment evaluation result of the video material to be evaluated is determined. The second matching degree between the theme keywords of the material and the audio text elements and visual elements is calculated to determine the material quality of the video material to be evaluated. This can achieve the purpose of video material quality evaluation based on segment analysis and overall analysis, and improve the comprehensiveness and accuracy of the identification results of low-quality video materials.
[0005] In a first aspect, embodiments of this application provide a method for determining low-quality video footage from mobile operators based on element object matching degree, the method comprising:
[0006] Obtain the video materials and themes associated with the target operator, determine the segment splitting length corresponding to the video materials to be evaluated, and split the video materials to be evaluated according to the segment splitting length to obtain multiple video segments;
[0007] Calculate the first matching degree between audio text elements and visual elements in each video segment, and determine the segment evaluation result of the video material to be evaluated based on the first matching degree corresponding to each video segment;
[0008] If the segment evaluation result is greater than the preset segment evaluation threshold, keywords are extracted from the material theme to obtain the material theme keywords, and the second matching degree between the material theme keywords and audio text elements and visual elements is calculated.
[0009] The matching degree of the second matching degree is compared with the matching degree threshold, and the quality of the video material to be evaluated is determined based on the comparison result.
[0010] Furthermore, there are multiple audio text elements and multiple visual elements in each video segment;
[0011] Calculate the first match degree between audio text elements and visual elements in each video segment, including:
[0012] Semantic recognition is performed on multiple audio text elements and multiple visual elements in each video segment, and audio vectors and visual vectors for each video segment are constructed based on the semantic recognition results.
[0013] The element relevance and element consistency of each video segment are determined based on the audio vector and visual vector, and the element relevance weight and element consistency weight corresponding to each video segment are determined.
[0014] The first matching degree between audio text elements and visual elements in each video segment is calculated based on element relevance weight, element consistency weight, element relevance, and element consistency.
[0015] Furthermore, the element relevance and element consistency of each video segment are determined based on the audio vectors and visual vectors, including:
[0016] Calculate the vector similarity between audio vectors and visual vectors to obtain the element correlation of each video segment;
[0017] When the element correlation is greater than the preset element correlation threshold, the first element type of each audio vector element in the audio vector and the second element type of each visual vector element in the visual vector are identified respectively.
[0018] The audio vector is mapped to the audio set range according to the first element type, and the visual vector is mapped to the visual set range according to the second element type. The overlap range between the audio set range and the visual set range is calculated, and the element consistency of each video segment is determined based on the overlap range.
[0019] Furthermore, the first element type includes either a detailed audio text element type or a regular audio text element type;
[0020] Determine the element-related weights and element-consistency weights corresponding to each video segment, including:
[0021] Determine the number of detailed audio text elements corresponding to the detailed element type and the number of regular audio text elements corresponding to the regular element type in each video segment, and calculate the ratio of the number of detailed audio text elements to the number of regular audio text elements.
[0022] The quantity ratio is matched with a preset element consistency weight lookup table. Based on the matching results, the element consistency weight corresponding to each video segment is determined, and the element related weight corresponding to each video segment is determined based on the element consistency weight.
[0023] Furthermore, the keywords for the materials include target keywords and category keywords;
[0024] Calculate the second match degree between the subject keywords of the material and the audio text elements and visual elements, including:
[0025] Identify the audio text keywords corresponding to each audio text element and the visual keywords corresponding to each visual element. Then, based on the material category keywords, filter the audio text keywords and visual keywords as candidate keywords to obtain candidate audio text keywords and candidate visual keywords.
[0026] The target keywords for each material are expanded according to the material category keywords. Based on the keyword expansion results, the candidate audio text keywords and candidate visual keywords are finally screened. The second evaluation result of the video material to be evaluated is determined according to the final keyword screening results.
[0027] Furthermore, based on the final keyword screening results, a second evaluation result is determined for the video materials to be evaluated, including:
[0028] The number of overlapping keywords is determined based on the number of final audio text keywords and the number of final visual keywords in the final keyword filtering results.
[0029] The difference between the number of target keywords and the number of overlapping keywords in the material is calculated to obtain the second evaluation result of the video material to be evaluated.
[0030] Furthermore, determine the segment length corresponding to the video material to be evaluated, including:
[0031] Identify multiple timestamps, video frame data, and audio data in the video material to be evaluated. Perform audio-visual alignment on the video frame data and audio data based on multiple timestamps, and perform semantic recognition on the audio data. Based on the semantic recognition results, determine the first video segmentation timestamp corresponding to the video material to be evaluated among the multiple timestamps.
[0032] The inter-frame difference between adjacent video frames in the video material to be evaluated is calculated based on the video frame data. The second video segmentation timestamp corresponding to the video material to be evaluated is determined from multiple timestamps based on the inter-frame difference.
[0033] By integrating the first video segmentation timestamp and the second video segmentation timestamp, the segment split length corresponding to the video material to be evaluated is obtained.
[0034] Secondly, embodiments of this application provide a system for identifying low-quality video footage from mobile operators based on element object matching degrees, the system comprising:
[0035] The video splitting module is used to obtain the video materials to be evaluated and the material themes associated with the target operator, determine the segment splitting length corresponding to the video materials to be evaluated, and split the video materials to be evaluated according to the segment splitting length to obtain multiple video segments;
[0036] The segment evaluation module is used to calculate the first matching degree between audio text elements and visual elements in each video segment, and to determine the segment evaluation result of the video material to be evaluated based on the first matching degree corresponding to each video segment.
[0037] The matching degree calculation module is used to extract keywords from the material theme when the segment evaluation result is greater than the preset segment evaluation threshold, obtain the material theme keywords, and calculate the second matching degree between the material theme keywords and audio text elements and visual elements;
[0038] The material quality assessment module is used to compare the matching degree of the second matching degree with the preset matching degree threshold, and determine the material quality of the video material to be evaluated based on the comparison result of the matching degree relationship.
[0039] Thirdly, embodiments of this application provide an electronic device including a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the method described in the first aspect.
[0040] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.
[0041] Fifthly, embodiments of this application also provide a computer program product comprising a computer program stored in a computer-readable storage medium, wherein at least one processor of the device reads from the computer-readable storage medium and executes the computer program, causing the device to perform the method described in the first aspect.
[0042] In this embodiment, the video material to be evaluated and the material theme associated with the target operator are obtained. The segment splitting length corresponding to the video material to be evaluated is determined, and the video material to be evaluated is split into multiple video segments according to the segment splitting length. The first matching degree between audio text elements and visual elements in each video segment is calculated, and the segment evaluation result of the video material to be evaluated is determined based on the first matching degree corresponding to each video segment. If the segment evaluation result is greater than a preset segment evaluation threshold, keywords are extracted from the material theme to obtain material theme keywords, and the second matching degree between the material theme keywords and audio text elements and visual elements is calculated. The matching degree relationship between the second matching degree and the preset matching degree threshold is compared, and the material quality of the video material to be evaluated is determined according to the matching degree relationship comparison result. The above-described method for identifying low-quality video materials from operators based on element-object matching solves the problems of incomplete and inaccurate identification results in existing technologies. By calculating the first matching degree between audio-text elements and visual elements in each video segment, the segment evaluation result of the video material to be evaluated is determined. The second matching degree between the material's theme keywords and audio-text elements and visual elements is calculated to determine the material quality of the video material to be evaluated. This achieves the goal of evaluating video material quality based on segment analysis and overall analysis, improving the comprehensiveness and accuracy of the identification results of low-quality video materials. Attached Figure Description
[0043] Figure 1 This is a flowchart of a method for determining low-quality video materials from operators based on element object matching degree, provided in an embodiment of this application;
[0044] Figure 2 This is a flowchart illustrating the calculation of the first matching degree between audio text elements and visual elements, provided in an embodiment of this application.
[0045] Figure 3 This is a schematic diagram of the vector element mapping range provided in the real-time example of this application;
[0046] Figure 4 This is a flowchart illustrating the calculation of the second matching degree between the subject keywords and the elements of the material, provided in an embodiment of this application.
[0047] Figure 5 This is a structural block diagram of a system for determining low-quality video materials for operators based on element object matching degree, provided in an embodiment of this application.
[0048] Figure 6 This is a structural block diagram of the electronic device provided in the embodiments of this application. Detailed Implementation
[0049] To make the objectives, technical solutions, and advantages of this application clearer, specific embodiments of this application are described in detail below with reference to the accompanying drawings. It is understood that the specific embodiments described herein are merely for explaining this application and not for limiting it. Furthermore, it should be noted that, for ease of description, only the parts relevant to this application are shown in the drawings, not all of them. Before discussing exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe operations (or steps) as sequential processes, many of these operations can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the operations can be rearranged. The process can be terminated when its operation is completed, but may also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.
[0050] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0051] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0052] First, this solution can be used in scenarios where video footage is used for product demonstration and promotion, especially by operators to accurately showcase and promote information about newly launched products. By calculating the first matching degree between audio-text elements and visual elements in each video segment, the evaluation result of the video footage segment to be evaluated is determined. The second matching degree between the theme keywords of the footage and the audio-text elements and visual elements is calculated to determine the quality of the video footage to be evaluated. This achieves the goal of evaluating video footage quality based on segment analysis and overall analysis, improving the comprehensiveness and accuracy of the results for identifying low-quality video footage. Based on the above use case, it is understood that the execution entity for each step in this solution can be a computer device. This computer device refers to any electronic device with data computing, processing, and storage capabilities, such as mobile phones, PCs (Personal Computers), tablets, and other terminal devices, or it can be a server, etc. This application embodiment does not limit this.
[0053] The following description, in conjunction with the accompanying drawings, details a method and system for determining low-quality video materials for mobile operators based on element object matching degree, through specific embodiments and application scenarios.
[0054] Figure 1 This is a flowchart illustrating a method for determining low-quality video footage from operators based on element object matching, as provided in an embodiment of this application. Figure 1 As shown, the specific steps include the following:
[0055] S101, obtain the video material to be evaluated and the material theme associated with the target operator, determine the segment splitting length corresponding to the video material to be evaluated, and split the video material to be evaluated according to the segment splitting length to obtain multiple video segments.
[0056] The target operator can be the operator that requests a quality assessment of the video footage. The video footage to be assessed can be a video used to introduce or promote products launched by the target operator. The footage theme can be text representing the main or core content of each video footage. The segment split length can be the segment length of the video footage to be assessed, divided into multiple video segments.
[0057] In one embodiment, video materials to be evaluated and material themes associated with the target operator can be obtained. Based on the preset correspondence between video material types and split lengths and the material type of the video materials to be evaluated, the segment split length corresponding to the video materials to be evaluated is determined, and the video materials to be evaluated are split into multiple video segments according to the segment split length.
[0058] In one embodiment, determining the segment split length corresponding to the video material to be evaluated includes: identifying multiple timestamps, video frame data, and audio data in the video material to be evaluated; performing audio-visual alignment on the video frame data and audio data based on the multiple timestamps; performing semantic recognition on the audio data; determining a first video segmentation timestamp corresponding to the video material to be evaluated from among the multiple timestamps based on the semantic recognition result; calculating the inter-frame difference degree between adjacent video frames in the video material to be evaluated based on the video frame data; determining a second video segmentation timestamp corresponding to the video material to be evaluated from among the multiple timestamps based on the inter-frame difference degree; and integrating the first video segmentation timestamp and the second video segmentation timestamp to obtain the segment split length corresponding to the video material to be evaluated.
[0059] The first video segmentation timestamp can represent the start or end timestamp of each sentence in the audio data. The inter-frame difference can represent the degree of change in images between two adjacent video frames. The second video segmentation timestamp can represent the start or end timestamp of the same action or related motion as a whole in the video frame data.
[0060] In one embodiment, multiple timestamps, video frame data, and audio data in the video material to be evaluated can be identified. The multiple timestamps can be located at corresponding positions in the video frame data and audio data, respectively. Audio-visual alignment is performed on the video frame data and audio data based on the alignment of identical timestamps among the multiple timestamps. Speech-to-text processing is performed on the audio data, and semantic recognition is performed on the converted audio text to determine the start or end position of each sentence in the audio text, and to determine the first video segmentation timestamp corresponding to the start or end position among the multiple timestamps. The inter-frame similarity between adjacent video frames in the video frame data can be calculated, and the inter-frame difference between adjacent video frames can be determined based on the inter-frame similarity. Adjacent video frames with an inter-frame difference greater than a preset inter-frame difference threshold are considered as video frames with different actions or movements, and the timestamp corresponding to any one of these two video frames is used as the second video segmentation timestamp. The video segmentation length corresponding to the first video segmentation timestamp is compared with the video segmentation length corresponding to the second video segmentation timestamp, and the larger video segmentation length is used as the segment segmentation length corresponding to the video material to be evaluated.
[0061] This solution determines the first video segmentation timestamp corresponding to the video material to be evaluated by semantic recognition of audio data, determines the second video segmentation timestamp corresponding to the video material to be evaluated based on the inter-frame difference between adjacent video frames, and integrates the first and second video segmentation timestamps to obtain the segment split length corresponding to the video material to be evaluated. This can achieve the purpose of video segmentation based on the correlation between video audio and video, and improve the rationality of the video segmentation length.
[0062] S102, calculate the first matching degree between audio text elements and visual elements in each video segment, and determine the segment evaluation result of the video material to be evaluated based on the first matching degree corresponding to each video segment.
[0063] The audio text elements can be the core information of the audio content of each video segment. The visual elements can be quantifiable or describable visual feature information extracted from the image frames of each video segment. The first matching degree can be a parameter representing the degree of consistency between the audio content and video content of each video segment. The segment evaluation result can be a parameter representing the overall matching status of all segments of the video material to be evaluated.
[0064] In one embodiment, the audio text elements and visual elements in each video segment can be quantified and formatted uniformly, and the similarity between the uniform audio text elements and visual elements can be calculated to obtain the first matching degree between the audio text elements and visual elements in each video segment. The average value of the first matching degree corresponding to each video segment is then calculated to obtain the segment evaluation result of the video material to be evaluated.
[0065] S103, if the segment evaluation result is greater than the preset segment evaluation threshold, extract keywords from the material theme to obtain material theme keywords, and calculate the second matching degree between the material theme keywords and audio text elements and visual elements.
[0066] The preset segment evaluation threshold can be a pre-set minimum value for the segment evaluation result under the condition that the audio and video matching of each segment in the video material to be evaluated is qualified. The second matching degree can be a parameter representing the degree of consistency between the overall audio and video content of the video material to be evaluated and the theme of the material.
[0067] In one embodiment, if the segment evaluation result is greater than a preset segment evaluation threshold, it indicates that the audio and video matching of each segment in the video material to be evaluated is qualified. At this time, it is also necessary to evaluate the overall audio and video matching degree of the video material with the material theme. Keywords can be extracted from the material theme using keyword extraction technology, and the similarity between the material theme keywords and audio text elements, as well as the similarity between the material theme keywords and visual elements, can be calculated. The minimum similarity is taken as the second matching degree.
[0068] S104. Compare the matching degree of the second matching degree with the matching degree threshold, and determine the quality of the video material to be evaluated based on the comparison result of the matching degree.
[0069] In one embodiment, a second matching degree can be compared with a preset matching degree threshold. If the second matching degree is greater than or equal to the preset matching degree threshold, the quality of the video material to be evaluated is determined to be qualified. If the second matching degree is less than the preset matching degree threshold, the quality of the video material to be evaluated is determined to be unqualified.
[0070] The technical solution provided in this application involves obtaining video materials and material themes associated with a target operator, determining the segment splitting length corresponding to the video materials to be evaluated, and splitting the video materials to be evaluated into multiple video segments according to the segment splitting length; calculating the first matching degree between audio text elements and visual elements in each video segment, and determining the segment evaluation result of the video materials to be evaluated based on the first matching degree corresponding to each video segment; if the segment evaluation result is greater than a preset segment evaluation threshold, extracting keywords from the material theme to obtain material theme keywords, and calculating the second matching degree between the material theme keywords and audio text elements and visual elements; comparing the matching degree relationship between the second matching degree and the preset matching degree threshold, and determining the material quality of the video materials to be evaluated based on the comparison result of the matching degree relationship. The above-described method for identifying low-quality video materials from operators based on element-object matching solves the problems of incomplete and inaccurate identification results in existing technologies. By calculating the first matching degree between audio-text elements and visual elements in each video segment, the segment evaluation result of the video material to be evaluated is determined. The second matching degree between the material's theme keywords and audio-text elements and visual elements is calculated to determine the material quality of the video material to be evaluated. This achieves the goal of evaluating video material quality based on segment analysis and overall analysis, improving the comprehensiveness and accuracy of the identification results of low-quality video materials.
[0071] Figure 2 This is a flowchart illustrating the calculation of the first matching degree between audio text elements and visual elements, provided in an embodiment of this application. For example... Figure 2 As shown, each video clip contains multiple audio text elements and multiple visual elements, specifically including the following steps:
[0072] S201, Semantic recognition is performed on multiple audio text elements and multiple visual elements in each video segment, and audio vectors and visual vectors of each video segment are constructed based on the semantic recognition results.
[0073] In one embodiment, semantic recognition can be performed on multiple audio text elements and multiple visual elements in each video segment. Audio text elements with the same meaning in the semantic recognition results are merged and used as the same vector element in the audio vector. All merged vector elements and other audio text elements with different meanings are integrated according to the vector format to obtain the audio vector of each video segment. Similarly, multiple visual elements are integrated to construct the visual vector of each video segment.
[0074] S202, determine the element relevance and element consistency of each video segment based on the audio vector and visual vector, and determine the element relevance weight and element consistency weight corresponding to each video segment.
[0075] Element relevance refers to the overall similarity between the meanings of audio and visual elements in each video segment. Element consistency refers to the degree of overlap between audio and visual elements with the same meaning in each video segment.
[0076] In one embodiment, the number of vector elements with the same or similar meaning in the audio and visual vectors can be identified, and the ratio of the number of vector elements with the same or similar meaning to the average number of all elements in the audio and visual vectors can be calculated to obtain the element relevance of each video segment. The number of elements with the same element value in the audio and visual vectors can be identified, and the ratio of the number of elements with the same element value to the average number of all elements can be calculated to obtain the element consistency. Based on a preset correspondence between element relevance and element relevance weights, the element relevance weight corresponding to each video segment is determined, and based on a preset correspondence between element consistency and element consistency weights, the element consistency weight corresponding to each video segment is determined.
[0077] In one embodiment, determining the element relevance and element consistency of each video segment based on audio vectors and visual vectors includes: calculating the vector similarity between audio vectors and visual vectors to obtain the element relevance of each video segment; if the element relevance is greater than a preset element relevance threshold, identifying the first element type of each audio vector element in the audio vector and the second element type of each visual vector element in the visual vector; mapping the audio vector to an audio set range based on the first element type, mapping the visual vector to a visual set range based on the second element type, calculating the overlap range between the audio set range and the visual set range, and determining the element consistency of each video segment based on the overlap range.
[0078] The preset element relevance threshold can be the minimum vector relevance when the overall meanings of the audio and visual vectors are similar. The first element type can be the element category corresponding to each audio vector element. The second element type can be the element category corresponding to each visual vector element.
[0079] In one embodiment, the cosine similarity between audio vectors and visual vectors can be calculated to obtain the element relevance of each video segment, and the element relevance can be compared with a preset element relevance threshold. If the element relevance is greater than the preset threshold, the first element type of each audio vector element and the second element type of each visual vector element are identified. Multiple element type distribution maps can be preset, and element types with the same first element type in the distribution maps are grouped together to obtain the audio set range mapped by the audio vectors. Similarly, element types with the same second element type in the distribution maps are grouped together to obtain the visual set range mapped by the visual vectors. The intersection of the audio set range and the visual set range is calculated to obtain the overlapping range, and the union of the audio set range and the visual set range is calculated to obtain the overall range. The proportion of the overlapping range in the overall range is calculated to obtain the element consistency of each video segment.
[0080] Figure 3 This is a schematic diagram of the vector element mapping range provided in the real-time example of this application. For example... Figure 3 As shown in the diagram, the distribution range includes multiple preset element types. This range contains most of the element types that might be used in the operator's video materials, including: brand slogans, product descriptions, brand prompts, 4G data traffic, 5G data traffic, brand theme songs, A1 operators, A2 operators, A3 operators, home scenes, service hall scenes, outdoor scenes, usage rules, broadband, calls, B2 style, character images, subtitles, brand logos, and character images. The area enclosed by the solid lines in the diagram represents the audio set range obtained by mapping audio vectors to the distribution map, and the area enclosed by the dashed lines represents the visual set range obtained by mapping visual vectors to the distribution map. The overlapping area between the dashed and solid lines is the intersection of the audio and visual set ranges, i.e., the overlapping range. Based on the number of elements in the union and intersection of the audio and visual set ranges, the proportion of the overlapping range in the overall range can be calculated, yielding the element consistency of each video segment.
[0081] This scheme calculates the vector similarity between audio and visual vectors to obtain the element relevance of each video segment. Based on the element type, the audio and visual vectors are mapped to the corresponding set ranges, and the overlap range is calculated to determine the element consistency of each video segment. This can improve the accuracy of the calculation results of element relevance and element consistency.
[0082] In one embodiment, the first element type includes a detailed audio text element type or a regular audio text element type; determining the element-related weight and element-consistency weight corresponding to each video segment includes: determining the number of detailed audio text elements corresponding to the detailed element type and the number of regular audio text elements corresponding to the regular element type in each video segment, and calculating the ratio of the number of detailed audio text elements to the number of regular audio text elements; matching the ratio with a preset element-consistency weight lookup table, determining the element-consistency weight corresponding to each video segment based on the matching result, and determining the element-related weight corresponding to each video segment based on the element-consistency weight.
[0083] Detailed audio text element types can be categories corresponding to elements that play a decisive role in the semantics of the audio text. Regular audio text element types can be categories corresponding to elements that make the semantics of the audio text complete.
[0084] In one embodiment, the number of detailed audio-text elements corresponding to detailed element types and the number of regular audio-text elements corresponding to regular element types in each video segment can be determined, and the ratio of the number of detailed audio-text elements to the number of regular audio-text elements in the same video segment can be calculated. This ratio is then matched against the range of ratios corresponding to the preset element consistency weights in a preset element consistency weight lookup table to obtain the element consistency weights corresponding to each video segment, and these element consistency weights are used as the element-related weights corresponding to each video segment.
[0085] This scheme improves the efficiency of calculating element-related weights by determining the number of detailed audio-text elements corresponding to detailed element types and the number of regular audio-text elements corresponding to regular element types in each video segment, calculating the ratio of the number of elements, and determining the element-related weights corresponding to each video segment based on the ratio of the number of elements.
[0086] S203, calculate the first matching degree between audio text elements and visual elements in each video segment based on element relevance weight, element consistency weight, element relevance degree and element consistency degree.
[0087] In one embodiment, the weighted sum of element relevance and element consistency in each video segment can be calculated based on element relevance weight, element consistency weight, element relevance, and element consistency to obtain the first matching degree between audio text elements and visual elements.
[0088] The technical solution provided in this application embodiment performs semantic recognition on multiple audio text elements and multiple visual elements in each video segment, constructs audio vectors and visual vectors for each video segment, determines the element relevance and element consistency of each video segment based on the audio vectors and visual vectors, determines the element relevance weight and element consistency weight corresponding to each video segment, and calculates the first matching degree between the audio text elements and visual elements in each video segment. This can improve the efficiency of calculating the first matching degree and the accuracy of the calculation results, which is beneficial to improving the accuracy of video segment quality assessment.
[0089] Figure 4 This is a flowchart illustrating the calculation of the second matching degree between the subject keywords and material elements, provided in an embodiment of this application. For example... Figure 4 As shown, the subject keywords of the materials include target keywords and category keywords, and the specific steps are as follows:
[0090] S401, determine the audio text keywords corresponding to each audio text element and the visual keywords corresponding to each visual element, and filter the audio text keywords and visual keywords based on the material category keywords to obtain candidate audio text keywords and candidate visual keywords.
[0091] Among them, the target keywords for the video material can be words that express the main theme of the video material. The category keywords for the video material can be words that express the direction or type of the video material. Candidate audio text keywords can be audio text keywords that are the same as the category keywords. Candidate visual keywords can be visual keywords that are the same as the category keywords.
[0092] In one embodiment, audio text keywords corresponding to each audio text element and visual keywords corresponding to each visual element can be determined based on the correspondence between elements and keywords. The consistency of material category keywords with audio text keywords and visual keywords can be verified. Consistent audio text keywords are used as candidate audio text keywords, and consistent visual keywords are used as candidate visual keywords.
[0093] S402, expand the target keywords of each material according to the material category keywords, and perform final keyword screening on the candidate audio text keywords and candidate visual keywords based on the keyword expansion results, and determine the second evaluation result of the video material to be evaluated based on the final keyword screening results.
[0094] In one embodiment, the target keywords of each material can be expanded using synonyms within the category of the material category to obtain keyword expansion results. These expanded results are then compared with candidate audio text keywords and candidate visual keywords. Candidate audio text keywords that match the expanded results are taken as final audio text keywords, and candidate visual keywords that match the expanded results are taken as final visual keywords. The sum of the final audio text keyword count and the final visual keyword count is calculated, along with the sum of the candidate audio text keyword count and the candidate visual keyword count. The ratio of the final sum to the candidate sum is then used to obtain the second evaluation result for the video material to be evaluated.
[0095] In one embodiment, determining the second evaluation result of the video material to be evaluated based on the final keyword screening results includes: determining the number of overlapping keywords based on the number of final audio text keywords and the number of final visual keywords in the final keyword screening results; calculating the difference between the number of target keywords and the number of overlapping keywords in the material to obtain the second evaluation result of the video material to be evaluated.
[0096] The number of overlapping keywords can be the minimum of the final audio text keywords and the final visual keywords.
[0097] In one embodiment, the number of final audio text keywords and the number of final visual keywords in the final keyword filtering results can be compared, the minimum number can be taken as the number of overlapping keywords, and the difference between the number of target keywords and the number of overlapping keywords can be calculated to obtain the second evaluation result of the video material to be evaluated.
[0098] This solution improves the efficiency of calculating the second evaluation result by determining the number of overlapping keywords and calculating the difference between the number of target keywords and the number of overlapping keywords in the material.
[0099] The technical solution provided in this application uses material category keywords to screen candidate keywords for audio text keywords and visual keywords respectively, obtaining candidate audio text keywords and candidate visual keywords. It expands the target keywords of each material according to the material category keywords, and performs final keyword screening on the candidate audio text keywords and candidate visual keywords based on the keyword expansion results. Finally, it determines the second evaluation result of the video material to be evaluated based on the final keyword screening results. This can improve the comprehensiveness and accuracy of the second evaluation result, which is conducive to improving the accuracy of the overall video material quality evaluation result.
[0100] Figure 5This is a structural block diagram of a system for determining low-quality video materials from operators based on element object matching degree, provided in an embodiment of this application. Figure 5 As shown, it specifically includes the following:
[0101] The video splitting module 501 is used to obtain the video material to be evaluated and the material theme associated with the target operator, determine the segment splitting length corresponding to the video material to be evaluated, and split the video material to be evaluated according to the segment splitting length to obtain multiple video segments.
[0102] The segment evaluation module 502 is used to calculate the first matching degree between audio text elements and visual elements in each video segment, and to determine the segment evaluation result of the video material to be evaluated based on the first matching degree corresponding to each video segment.
[0103] The matching degree calculation module 503 is used to extract keywords from the material theme when the segment evaluation result is greater than the preset segment evaluation threshold, obtain the material theme keywords, and calculate the second matching degree between the material theme keywords and the audio text elements and visual elements;
[0104] The material quality assessment module 504 is used to compare the matching degree of the second matching degree with the preset matching degree threshold, and to determine the material quality of the video material to be evaluated based on the comparison result of the matching degree relationship.
[0105] Furthermore, there are multiple audio text elements and multiple visual elements in each video segment;
[0106] Fragment evaluation module 502 is specifically used for:
[0107] Semantic recognition is performed on multiple audio text elements and multiple visual elements in each video segment, and audio vectors and visual vectors for each video segment are constructed based on the semantic recognition results.
[0108] The element relevance and element consistency of each video segment are determined based on the audio vector and visual vector, and the element relevance weight and element consistency weight corresponding to each video segment are determined.
[0109] The first matching degree between audio text elements and visual elements in each video segment is calculated based on element relevance weight, element consistency weight, element relevance, and element consistency.
[0110] Furthermore, the fragment evaluation module 502 is specifically used for:
[0111] Calculate the vector similarity between audio vectors and visual vectors to obtain the element correlation of each video segment;
[0112] When the element correlation is greater than the preset element correlation threshold, the first element type of each audio vector element in the audio vector and the second element type of each visual vector element in the visual vector are identified respectively.
[0113] The audio vector is mapped to the audio set range according to the first element type, and the visual vector is mapped to the visual set range according to the second element type. The overlap range between the audio set range and the visual set range is calculated, and the element consistency of each video segment is determined based on the overlap range.
[0114] Furthermore, the first element type includes either a detailed audio text element type or a regular audio text element type;
[0115] Fragment evaluation module 502 is specifically used for:
[0116] Determine the number of detailed audio text elements corresponding to the detailed element type and the number of regular audio text elements corresponding to the regular element type in each video segment, and calculate the ratio of the number of detailed audio text elements to the number of regular audio text elements.
[0117] The quantity ratio is matched with a preset element consistency weight lookup table. Based on the matching results, the element consistency weight corresponding to each video segment is determined, and the element related weight corresponding to each video segment is determined based on the element consistency weight.
[0118] Furthermore, the keywords for the materials include target keywords and category keywords;
[0119] Matching degree calculation module 503 is specifically used for:
[0120] Identify the audio text keywords corresponding to each audio text element and the visual keywords corresponding to each visual element. Then, based on the material category keywords, filter the audio text keywords and visual keywords as candidate keywords to obtain candidate audio text keywords and candidate visual keywords.
[0121] The target keywords for each material are expanded according to the material category keywords. Based on the keyword expansion results, the candidate audio text keywords and candidate visual keywords are finally screened. The second evaluation result of the video material to be evaluated is determined according to the final keyword screening results.
[0122] Furthermore, the matching degree calculation module 503 is specifically used for:
[0123] The number of overlapping keywords is determined based on the number of final audio text keywords and the number of final visual keywords in the final keyword filtering results.
[0124] The difference between the number of target keywords and the number of overlapping keywords in the material is calculated to obtain the second evaluation result of the video material to be evaluated.
[0125] Furthermore, the video splitting module 501 is specifically used for:
[0126] Identify multiple timestamps, video frame data, and audio data in the video material to be evaluated. Perform audio-visual alignment on the video frame data and audio data based on multiple timestamps, and perform semantic recognition on the audio data. Based on the semantic recognition results, determine the first video segmentation timestamp corresponding to the video material to be evaluated among the multiple timestamps.
[0127] The inter-frame difference between adjacent video frames in the video material to be evaluated is calculated based on the video frame data. The second video segmentation timestamp corresponding to the video material to be evaluated is determined from multiple timestamps based on the inter-frame difference.
[0128] By integrating the first video segmentation timestamp and the second video segmentation timestamp, the segment split length corresponding to the video material to be evaluated is obtained.
[0129] The technical solution provided in this application includes a video segmentation module, which is used to acquire video materials to be evaluated and material themes associated with a target operator, determine the segment segmentation length corresponding to the video materials to be evaluated, and perform video segmentation on the video materials to be evaluated according to the segment segmentation length to obtain multiple video segments; a segment evaluation module, which is used to calculate the first matching degree between audio text elements and visual elements in each video segment, and determine the segment evaluation result of the video materials to be evaluated based on the first matching degree corresponding to each video segment; a matching degree calculation module, which is used to extract keywords from the material theme when the segment evaluation result is greater than a preset segment evaluation threshold, obtain material theme keywords, and calculate the second matching degree between the material theme keywords and audio text elements and visual elements; and a material quality evaluation module, which is used to compare the matching degree relationship between the second matching degree and the preset matching degree threshold, and determine the material quality of the video materials to be evaluated based on the matching degree relationship comparison result. The aforementioned system for identifying low-quality video materials from operators based on element-object matching solves the problems of incomplete and inaccurate identification results in existing technologies. By calculating the first matching degree between audio-text elements and visual elements in each video segment, the system determines the segment evaluation result of the video material to be evaluated. By calculating the second matching degree between the material's theme keywords and audio-text elements and visual elements, the system determines the material quality of the video material to be evaluated. This achieves the goal of evaluating video material quality based on segment analysis and overall analysis, improving the comprehensiveness and accuracy of the identification results for low-quality video materials.
[0130] The low-quality video material identification system for operators based on element object matching degree in this application embodiment can be configured in a device, or in a component, integrated circuit, or chip in a terminal. The system can be configured in mobile electronic devices or non-mobile electronic devices. For example, mobile electronic devices can be mobile phones, tablets, laptops, PDAs, in-vehicle electronic devices, wearable devices, ultra-mobile personal computers (UMPCs), netbooks, or personal digital assistants (PDAs), etc., while non-mobile electronic devices can be servers, network attached storage (NAS), personal computers (PCs), televisions (TVs), ATMs, or self-service machines, etc. This application embodiment does not specifically limit the scope of the system.
[0131] The low-quality video material identification system for operators based on element object matching degree in this application embodiment can be an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit it.
[0132] The system for determining low-quality video materials from operators based on element object matching degree provided in this application embodiment can realize the various processes implemented in the above method embodiments. To avoid repetition, it will not be described again here.
[0133] like Figure 6 As shown, this application embodiment also provides an electronic device 600, including a processor 601, a memory 602, and a program or instructions stored in the memory 602 and executable on the processor 601. When the program or instructions are executed by the processor 601, they implement the various processes of the above-described embodiment of a method for determining low-quality video materials of operators based on element object matching degree, and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0134] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.
[0135] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described method for determining low-quality video materials of operators based on element object matching degree, and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0136] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0137] This application also provides a program product including program code. When the program product is run on a computer device, the program code causes the computer device to perform the steps of the methods described above according to various exemplary embodiments of this application. For example, the computer device can execute a method for determining low-quality video footage from a carrier based on element object matching degree, as described in an embodiment of this application. The program product can be implemented using any combination of one or more readable media.
[0138] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element. Furthermore, it should be noted that the scope of the methods and systems in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0139] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0140] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
[0141] The above description is merely a preferred embodiment and the technical principles employed in this application. This application is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions that can be made by those skilled in the art will not depart from the scope of protection of this application. Therefore, although this application has been described in detail through the above embodiments, this application is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of this application, the scope of which is determined by the scope of the claims.
Claims
1. A method for identifying low-quality video footage from mobile operators based on element object matching, characterized in that, The method includes: Obtain the video material to be evaluated and the material theme associated with the target operator, determine the segment splitting length corresponding to the video material to be evaluated, and split the video material to be evaluated into multiple video segments according to the segment splitting length; Calculating the first matching degree between multiple audio text elements and multiple visual elements in each video segment includes: performing semantic recognition on multiple audio text elements and multiple visual elements in each video segment respectively; constructing audio vectors and visual vectors for each video segment based on the semantic recognition results; calculating the vector similarity between the audio vectors and the visual vectors to obtain the element relevance of each video segment; if the element relevance is greater than a preset element relevance threshold, identifying the first element type of each audio vector element in the audio vector and the second element type of each visual vector element in the visual vector; and determining the element relevance based on the first element type. The method maps the audio vector to an audio set range, maps the visual vector to a visual set range according to the second element type, calculates the overlap range between the audio set range and the visual set range, determines the element consistency degree of each video segment based on the overlap range, and determines the element relevance weight and element consistency weight corresponding to each video segment. Based on the element relevance weight, the element consistency weight, the element relevance degree and the element consistency degree, the first matching degree between the audio text elements and visual elements in each video segment is calculated, and the segment evaluation result of the video material to be evaluated is determined based on the first matching degree corresponding to each video segment. If the segment evaluation result is greater than the preset segment evaluation threshold, keywords are extracted from the material theme to obtain material theme keywords, and the second matching degree between the material theme keywords and the audio text element and the visual element is calculated; The matching degree of the second matching degree is compared with the matching degree of the preset matching degree threshold, and the quality of the video material to be evaluated is determined based on the comparison result of the matching degree.
2. The method for determining low-quality video materials from operators based on element object matching degree according to claim 1, characterized in that, The first element type includes either a detailed audio text element type or a regular audio text element type; The determination of the element-related weights and element-consistency weights corresponding to each of the video segments includes: Determine the number of detailed audio text elements corresponding to the detailed audio text element type and the number of regular audio text elements corresponding to the regular audio text element type in each video segment, and calculate the ratio of the number of detailed audio text elements to the number of regular audio text elements; The quantity ratio is matched with a preset element consistency weight lookup table. The element consistency weight corresponding to each video segment is determined based on the matching result. The element related weight corresponding to each video segment is determined based on the element consistency weight.
3. The method for determining low-quality video materials from operators based on element object matching degree according to claim 1, characterized in that, The subject keywords of the materials include target keywords and category keywords; The calculation of the second matching degree between the material's theme keywords and the audio text elements and the visual elements includes: Determine the audio text keywords corresponding to each of the audio text elements and the visual keywords corresponding to each of the visual elements, and filter the audio text keywords and visual keywords based on the material category keywords to obtain candidate audio text keywords and candidate visual keywords; The target keywords of each material are expanded according to the keywords of the material category. Based on the keyword expansion results, the candidate audio text keywords and the candidate visual keywords are finally screened. The second evaluation result of the video material to be evaluated is determined according to the final keyword screening result.
4. The method for determining low-quality video materials from operators based on element object matching degree according to claim 3, characterized in that, The second evaluation result for determining the video material to be evaluated based on the final keyword filtering results includes: The number of overlapping keywords is determined based on the number of final audio text keywords and the number of final visual keywords in the final keyword filtering results. The difference between the number of target keywords in the material and the number of overlapping keywords is calculated to obtain the second evaluation result of the video material to be evaluated.
5. The method for determining low-quality video materials from operators based on element object matching degree according to claim 1, characterized in that, Determining the segment split length corresponding to the video material to be evaluated includes: Multiple timestamps, video frame data, and audio data in the video material to be evaluated are identified. Based on the multiple timestamps, the video frame data and audio data are aligned for audio and video, and semantic recognition is performed on the audio data. Based on the semantic recognition results, the first video segmentation timestamp corresponding to the video material to be evaluated is determined from among the multiple timestamps. Based on the video frame data, the inter-frame difference degree between adjacent video frames in the video material to be evaluated is calculated, and the second video segmentation timestamp corresponding to the video material to be evaluated is determined from the plurality of timestamps according to the inter-frame difference degree. By integrating the first video segmentation timestamp and the second video segmentation timestamp, the segment splitting length corresponding to the video material to be evaluated is obtained.
6. A system for identifying low-quality video materials from telecom operators based on element-object matching, characterized in that, The system includes: The video splitting module is used to obtain the video material to be evaluated and the material theme associated with the target operator, determine the segment splitting length corresponding to the video material to be evaluated, and split the video material to be evaluated into multiple video segments according to the segment splitting length. The segment evaluation module is used to calculate the first matching degree between multiple audio text elements and multiple visual elements in each video segment. This includes: performing semantic recognition on the multiple audio text elements and multiple visual elements in each video segment; constructing audio vectors and visual vectors for each video segment based on the semantic recognition results; calculating the vector similarity between the audio vectors and the visual vectors to obtain the element relevance of each video segment; and, if the element relevance is greater than a preset element relevance threshold, identifying the first element type of each audio vector element in the audio vector and the second element type of each visual vector element in the visual vector, according to the... The first element type maps the audio vector to an audio set range, the second element type maps the visual vector to a visual set range, and calculates the overlap range between the audio set range and the visual set range. Based on the overlap range, the element consistency degree of each video segment is determined, and the element relevance weight and element consistency weight corresponding to each video segment are determined. Based on the element relevance weight, the element consistency weight, the element relevance degree, and the element consistency degree, the first matching degree between the audio text elements and visual elements in each video segment is calculated, and the segment evaluation result of the video material to be evaluated is determined based on the first matching degree corresponding to each video segment. The matching degree calculation module is used to extract keywords from the material theme when the segment evaluation result is greater than a preset segment evaluation threshold, obtain material theme keywords, and calculate the second matching degree between the material theme keywords and the audio text element and the visual element; The material quality assessment module is used to compare the matching degree of the second matching degree with the matching degree threshold, and determine the material quality of the video material to be assessed based on the comparison result of the matching degree.
7. An electronic device, characterized in that, The method includes a processor, a memory, and a program or instructions stored in the memory and running on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the method for determining low-quality video material of an operator based on element object matching degree as described in any one of claims 1-5.
8. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the operator low-quality video material determination method based on element object matching degree as described in any one of claims 1-5.
Citation Information
Patent Citations
Video quality assessment method and device and storage medium
CN113822876A
Video quality detection method, equipment, storage medium and device
CN114760460A