Video file film and television information matching method and device, terminal and storage medium
By identifying the target keywords of the video file and querying the integrated media data set from multiple media libraries, and combining the metadata of the video file for weight matching, the problem of inaccurate matching of video files in the prior art is solved, and higher matching accuracy and user experience are achieved.
Patent Information
- Application Number
- CN202411904303.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-05-30
AI Technical Summary
The existing video file film and television information matching methods are difficult to accurately match the film and television information corresponding to the video file when facing complex file naming rules and multi-layer directory structures.
The target keyword is identified by obtaining the video file information of the video file in the target storage device and using the pre-trained target recognition model. According to the identification results, the integrated media data set is obtained from multiple media databases, and the weight matching calculation is performed with the metadata of the video file to obtain the optimal integrated media data corresponding to the video file.
It improves the accuracy of video file film and television information matching, enhances user experience, and can accurately match film and television information in complex file naming rules and multi-layer directory structures.
Smart Images

Figure CN120067392A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing and information technology, and in particular, to a method, device, terminal, and storage medium for matching film and television information of video files. Background Art
[0002] With the development of Internet technology, OTT (Over-the-Top) set-top boxes and smart TVs have gradually become the mainstream devices for home entertainment. Users manage and operate video files by connecting external storage devices or accessing remote file servers. However, existing methods for matching film and television information of video files often have difficulty accurately matching the corresponding film and television information of video files when faced with complex file naming rules and multi-level directory structures.
[0003] Therefore, how to accurately match the most suitable film and television information for video files in a storage device from multiple media libraries is an urgent problem to be solved. Summary of the Invention
[0004] Embodiments of the present invention provide a method, device, terminal, and storage medium for matching film and television information of video files to solve the technical problem of low accuracy in matching film and television information of existing video files.
[0005] In a first aspect, a method for matching film and television information of video files is provided, including: Obtain video file information of a video file in a target storage device; Input the video file information into a pre-trained target recognition model to obtain a recognition result output by the target recognition model; When the recognition result is that a target keyword is recognized, query a fusion media data set from multiple media libraries according to the target keyword; Extract metadata corresponding to the video file from the target storage device; Perform weighted matching calculation on each fusion media data in the fusion media data set with the metadata in sequence to obtain the optimal fusion media data corresponding to the video file, and store the optimal fusion media data in a database.
[0006] In an embodiment, the pre-trained target recognition model includes a pre-trained video recognition model and a video collection recognition model; The step of inputting the video file information into a pre-trained target recognition model to obtain a recognition result output by the target recognition model includes: Determine the category corresponding to the video file according to the video file information; When the category is the first category, input the video file information into the video recognition model to obtain a first recognition result output by the video recognition model; When the category is the second category, input the video file information into the video collection recognition model to obtain a second recognition result output by the video collection recognition model.
[0007] In one embodiment, the target keyword includes a first target keyword; When the recognition result is that the target keyword is recognized, querying a fused media asset dataset from multiple media asset libraries according to the target keyword includes: When the recognition result is the first recognition result, determine whether the first recognition result recognizes the first target keyword; If the first target keyword is recognized, use the first target keyword to query all first media asset data from multiple media asset libraries respectively; Use multi-source data fusion technology to fuse all the first media asset data to obtain the fused media asset dataset.
[0008] In one embodiment, the target keyword includes a second target keyword; When the recognition result is that the target keyword is recognized, querying a fused media asset dataset from multiple media asset libraries according to the target keyword includes: When the recognition result is the second recognition result, determine whether the second recognition result recognizes the second target keyword; If the second target keyword is recognized, use the second target keyword to query all second media asset data from multiple media asset libraries respectively; Use multi-source data fusion technology to fuse all the second media asset data to obtain the fused media asset dataset.
[0009] In one embodiment, the step of performing weighted matching calculation on each fused media asset data in the fused media asset dataset with the metadata in sequence to obtain the optimal fused media asset data corresponding to the video file includes: Perform weighted matching calculation on each fused media asset data in the fused media asset dataset with the metadata in sequence to obtain a comprehensive matching score corresponding to each fused media asset data; Based on the comprehensive matching scores, sort the fused media asset data in the fused media asset dataset to obtain a sorting result; Select the fused media asset data with the highest comprehensive matching score from the sorting result as the optimal fused media asset data corresponding to the video file.
[0010] In one embodiment, the step of sequentially performing weight matching calculation on each piece of fused media asset data in the fused media asset dataset with the metadata to obtain a comprehensive matching score corresponding to each piece of fused media asset data includes: According to the preset priorities of the metadata, the metadata is divided into first-level metadata and second-level metadata, and corresponding matching weights are set for the first-level metadata and the second-level metadata respectively; For each piece of fused media asset data in the fused media asset dataset, the following steps are performed: Calculate a first matching degree score between the fused media asset data and the first-level metadata; Calculate a second matching degree score between the fused media asset data and the second-level metadata; According to the first matching degree score and the matching weight corresponding to the first-level metadata, obtain a first weighted score; According to the second matching degree score and the matching weight corresponding to the second-level metadata, obtain a second weighted score; According to the first weighted score and the second weighted score, obtain the comprehensive matching score of the fused media asset data.
[0011] In one embodiment, the video file information includes the file name and / or path information corresponding to the video file, and the method further includes: When the recognition result is that the target keyword is not recognized, use a target algorithm to extract a target extraction keyword from the file name and / or path information; According to the target extraction keyword, query a fused media asset dataset from the multiple media asset libraries.
[0012] In a second aspect, a film and television information matching device for a video file is provided, including: A first acquisition module, configured to acquire video file information of a video file in a target storage device; An input module, configured to input the video file information into a pre-trained target recognition model to obtain a recognition result output by the target recognition model; A second acquisition module, configured to, when the recognition result is that a target keyword is recognized, query a fused media asset dataset from multiple media asset libraries according to the target keyword; An extraction module, configured to extract metadata corresponding to the video file from the target storage device; A storage module, configured to sequentially perform weight matching calculation on each piece of fused media asset data in the fused media asset dataset with the metadata to obtain the optimal fused media asset data corresponding to the video file, and store the fused media asset data corresponding to the video file in a database.
[0013] In a third aspect, an intelligent terminal is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for matching film and television information of a video file described in the first aspect above is implemented.
[0014] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for matching film and television information of a video file described in the first aspect above is implemented.
[0015] In a solution implemented by the method, device, terminal, and storage medium for matching film and television information of a video file described above, by obtaining video file information and using a target recognition model to identify target keywords, content related to the video file can be initially screened out, improving the pertinence of matching. Then, the video file metadata is extracted and weight matching calculations are performed with each piece of fusion media data in the fusion media dataset obtained based on the target keywords. This weight-based matching method can comprehensively consider various factors to evaluate the adaptability of the data, thereby accurately locating the optimal fusion media data (film and television information) and storing it. Overall, this method can effectively improve the accuracy of matching film and television information of video files and enhance the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0017] Figure 1 is a schematic diagram of an application environment of the method for matching film and television information of a video file in an embodiment of the present invention; Figure 2 is a flowchart of the method for matching film and television information of a video file in an embodiment of the present invention; Figure 3 is another flowchart of the method for matching film and television information of a video file in an embodiment of the present invention; Figure 4 is another flowchart of the method for matching film and television information of a video file in an embodiment of the present invention; Figure 5 is another flowchart of the method for matching film and television information of a video file in an embodiment of the present invention; Figure 6 is a schematic diagram of the device for matching film and television information of a video file in an embodiment of the present invention; Figure 7It is a schematic diagram of a smart terminal in an embodiment of the present invention. Detailed implementation manners
[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0019] The method for matching film and television information of a video file provided by an embodiment of the present invention can be applied to an application environment as Figure 1 shown. Specifically, the method for matching film and television information of a video file is applied in a system for matching film and television information of a video file, and the system for matching film and television information of a video file includes a smart terminal, a target storage device, and a display device as Figure 1 shown. Among them, the smart terminal includes, but is not limited to, various set-top boxes, smart TVs, personal computers, laptop computers, smart phones, and tablet computers. The target storage device includes, but is not limited to, a server and a mobile hard disk, and the display device includes, but is not limited to, various televisions or monitors. Next, the method for matching film and television information of a consistent video file provided by the present invention will be introduced in detail according to specific embodiments.
[0020] In a first aspect, as Figure 2 shown, a method for matching film and television information of a video file is provided. Taking the smart terminal in Figure 1 as an example for illustration, the method includes the following steps: S10. Obtain the video file information of the video file in the target storage device.
[0021] In this embodiment, the target storage device includes, but is not limited to, a mobile hard disk, a USB flash drive, a network attached storage (NAS), and a remote file server, etc.
[0022] The video file information includes the file name and / or path information of the video file, etc.
[0023] As an example, the video file can be scanned according to a specified target directory (for example, a specified folder) to obtain the video file information of the video file from the target storage device, providing the necessary data support for further accurately matching film and television information in the future.
[0024] For example, when the target storage device is an external hard drive, assuming that a target directory named "Video Resources" is specified, after successfully connecting to the external hard drive through the USB interface, the file scanning tool built into the smart terminal will start from the "Video Resources" directory and use the depth-first traversal algorithm to sequentially view each file and subdirectory under this directory. Once a video file is encountered, such as a file named "Interstellar Blu-ray Edition.MP4", the video file information corresponding to the video file is extracted. For example, the file name (such as Interstellar Blu-ray Edition.MP4) and the path where it is located (such as Video Resources / Science Fiction Movies / Interstellar).
[0025] S20. Input the video file information into a pre-trained target recognition model to obtain the recognition result output by the target recognition model.
[0026] In this embodiment, the recognition result includes recognizing the target keyword or not recognizing the target keyword.
[0027] As an example, after obtaining the video file information, the video file information will be input into a pre-trained target recognition model, thereby obtaining the recognition result output by the target recognition model. Among them, the pre-trained target recognition model is pre-trained based on the pre-set file name and path of the video file, and can automatically extract the target keyword from the file name and path of the video file. For example, target keywords such as the movie name and clarity are found from the file name, or the category corresponding to the video file is extracted from the path information, such as TV series, anime, or movie, and finally the recognition result is output.
[0028] For example, the video file information of a video file being processed is "E: / TV Series / Game of Thrones / Game of Thrones.2011.S01E01.Summer More Resources - XH1080.com.mkv". The video file information of this video file includes the file name "Game of Thrones.2011.S01E01.Summer More Resources - XH1080.com.mkv", the file path "E: / TV Series / Game of Thrones / ", etc. Further, the obtained video file information is input into a pre-trained target recognition model; by analyzing the input video file information, the target recognition model outputs corresponding target keywords, such as target keywords like Game of Thrones, TV series, 2011, S01E01, and 1080. Among them, in S01E01, "S" represents "Season", "01" indicates the first season; "E" represents "Episode", and "01" means the first episode of this season. It should be understood that the above is only an example and does not constitute a limitation to the present invention.
[0029] S30. When the recognition result is that a target keyword is recognized, according to the target keyword, query a fused media asset dataset from multiple media asset libraries.
[0030] In this embodiment, the target keyword may include one or more, and no limitation is made here.
[0031] The multiple media asset libraries may include but are not limited to The Movie Database (TMDB), Internet Movie Database (IMDB), and Douban, etc. It should be understood that the number of the multiple media asset libraries may be two or more, and no limitation is made here either.
[0032] As an example, after the target keyword is recognized, further perform data queries in multiple media asset libraries respectively through the obtained one or more target keywords to obtain all media asset data related to the video file, and fuse the obtained all media asset data using multi-source data fusion technology to obtain multiple fused media asset data, so as to form a complete fused media asset dataset from the obtained multiple fused media asset data. Among them, the fused media asset data includes but is not limited to basic information (such as movie name, director, actor, release year, etc.) and other information (such as user rating, film synopsis, poster image, etc.), and no limitation is made here.
[0033] S40. Extract the metadata corresponding to the video file from the target storage device.
[0034] In this embodiment, the metadata includes, but is not limited to, data such as the release year corresponding to the video file, video type, director, actors, country, clarity, and / or audio-video type.
[0035] As an example, a multimedia parsing tool (such as FFmpeg, MediaInfo) can be used, or the metadata corresponding to the video file can be extracted from the target storage device by accessing the embedded data of the video file itself (such as the metadata fields in the video container), which is not limited here.
[0036] S50. Sequentially perform weighted matching calculations on each piece of integrated media asset data in the integrated media asset data set with the metadata to obtain the optimal integrated media asset data corresponding to the video file, and store the optimal integrated media asset data in the database.
[0037] In this embodiment, the integrated media asset data set may include one or more pieces of integrated media asset data. If there are multiple pieces of integrated media asset data, it indicates that integrated media asset data corresponding to multiple versions of the video file are retrieved according to the target keyword. For example, multiple versions of the same movie, such as "Spider-Man 1", "Spider-Man 2", "Spider-Man 3", etc.
[0038] As an example, after obtaining the integrated media asset data set, first determine the number of pieces of integrated media asset data in the integrated media asset data set; if there is one, the integrated media asset data can be directly used as the optimal integrated media asset data to further improve the matching efficiency. Alternatively, the one piece of integrated media asset data can also be subjected to weighted matching calculation with the metadata to determine whether it is the integrated media asset data (film and television information) corresponding to the video file, thereby improving the accuracy of film and television information matching. If not, the user can be notified of incorrect information matching for timely correction.
[0039] If there are multiple pieces, each piece of integrated media asset data is sequentially subjected to weighted matching calculation with the metadata corresponding to the video file to obtain the optimal integrated media asset data corresponding to the video file, and the optimal integrated media asset data is stored in the database to ensure that the finally selected film and television information is the most matching version of the video file content, thus greatly improving the accuracy of matching.
[0040] For example, assume that the video file "SpiderMan.mp4" is being processed, and this video file is one of the "Spider-Man" series of films. The video file processing system obtains a fused media dataset from multiple media repositories (such as TMDB, IMDB, Douban), which contains multiple versions of fused media data. For example, "Spider-Man 1" (released in 2002, directed by Sam Raimi, starring Tobey Maguire, rating 8.2 / 10), "Spider-Man 2" (released in 2004, directed by Sam Raimi, starring Tobey Maguire, rating 7.9 / 10), and "Spider-Man 3" (released in 2007, directed by Sam Raimi, starring Tobey Maguire, rating 6.5 / 10). At the same time, the metadata extracted from the video file includes the release year 2002, director Sam Raimi, and star Tobey Maguire. Based on this metadata, the video file processing system performs a weighted matching calculation with each fused media data, and finds that "Spider-Man 1" exactly matches the video file in all key metadata. Therefore, it is selected as the optimal fused media data and stored in the database. The stored data includes information such as the film name "Spider-Man 1", director Sam Raimi, star Tobey Maguire, release year 2002, and rating 8.2 / 10.
[0041] Through the above method, even if the fused media dataset contains information about multiple versions of a film, it is still possible to accurately match the version that is most relevant to the content of the video file, thereby greatly improving the accuracy and reliability of the matching results.
[0042] It should be noted that the above is only described as an example and does not constitute a limitation to the present invention.
[0043] In summary, a solution implemented in an embodiment of the present application can initially screen out content related to a video file by obtaining video file information and using a target recognition model to identify target keywords, improving the pertinence of the matching; then extract the video file metadata and perform a weighted matching calculation with each fused media data in the fused media dataset obtained based on the target keywords. This weighted matching method can comprehensively consider various factors to measure the adaptability of the data, thereby accurately locating the optimal fused media data (film and television information) and storing it. Overall, this method can effectively improve the accuracy of film and television information matching for video files and improve the user experience.
[0044] In one embodiment, as Figure 3 shown, that is, in step S20, the pre-trained target recognition model includes a pre-trained video recognition model and a video collection recognition model; That is, in the step of inputting the video file information into a pre-trained target recognition model to obtain the recognition result output by the target recognition model, the following steps are included: S21. Determine the category corresponding to the video file according to the video file information; S22. When the category is the first category, input the video file information into the video recognition model to obtain the first recognition result output by the video recognition model; S23. When the category is the second category, input the video file information into the video collection recognition model to obtain the second recognition result output by the video collection recognition model.
[0045] In this embodiment, the target recognition model consists of two parts: a video recognition model and a video collection recognition model. According to the category of the video file, the video file processing system decides which model to input the video file information into for processing to obtain the corresponding recognition result.
[0046] Among them, the first category refers to an independent single film, such as a movie, etc.; the second category refers to a series of films, usually a video collection, such as a TV drama series, a movie series, and an anime collection, etc.
[0047] As an example, the category to which the video file belongs can be determined according to the video file information (such as the file name and path); when the video file belongs to the first category (for example, an independent movie file), the video file information is input into the video recognition model to obtain the first recognition result; when the video file belongs to the second category (for example, an episode in a TV drama series or a movie series), the video file information is input into the video collection recognition model to obtain the second recognition result.
[0048] Through this process, different recognition models can be used respectively according to the category of the video file to obtain the most accurate recognition result, which helps to process different types of video files, ensures that the most suitable recognition strategy is adopted for independent films and video collections, and thus improves the accuracy and efficiency of recognition.
[0049] It should be noted that the specific content included in the above first recognition result and second recognition result can refer to the description of the recognition result in step S20. To avoid repetition, it will not be elaborated here. And when it is recognized that the video file is of the second category, the matching rule of this time is recorded, and the video files in the same directory will not need to judge the category and will be directly processed according to the above same recognition model, so as to improve the processing efficiency.
[0050] In one embodiment, as Figure 3 shown, that is, in step S30, the target keyword includes the first target keyword; That is, when the recognition result is that the target keyword is recognized, the steps for querying the fused media asset dataset from multiple media asset libraries according to the target keyword include the following: S31A. When the recognition result is the first recognition result, determine whether the first recognition result recognizes the first target keyword; S32A. If the first target keyword is recognized, use the first target keyword to query all first media asset data from multiple media asset libraries respectively; S33A. Use multi-source data fusion technology to fuse all the first media asset data to obtain the fused media asset dataset.
[0051] As an example, when the first recognition result is obtained, first determine whether the first recognition result recognizes the first target keyword, such as the movie name, etc.; if the target keyword is recognized, use the first target keyword to initiate a query to multiple external media asset libraries (such as TMDB, IMDB, Douban, etc.) to obtain all the first media asset data related to the video file, and then use multi-source data fusion technology to fuse all the first media asset data obtained by the query to obtain multiple fused media asset data, and further combine the multiple fused media asset data to obtain the fused media asset dataset.
[0052] It should be understood that the target keyword here includes the first target keyword, which can be understood as the target keyword being the first target keyword. Here, it is just because different models are used for recognition and the obtained target keywords are also different, so a distinction is made.
[0053] In an embodiment, as Figure 3 shown, that is, in step S30, the target keyword includes the second target keyword; That is, when the recognition result is that the target keyword is recognized, the steps for querying the fused media asset dataset from multiple media asset libraries according to the target keyword further include the following: S31B. When the recognition result is the second recognition result, determine whether the second recognition result recognizes the second target keyword; S32B. If the second target keyword is recognized, use the second target keyword to query all second media asset data from multiple media asset libraries respectively; S33B. Use multi-source data fusion technology to fuse all the second media asset data to obtain the fused media asset dataset.
[0054] As an example, when the second recognition result is obtained, it is first determined whether the recognition result includes a second target keyword. For example, when the video file belongs to a certain series (such as a TV drama or movie series), the second target keyword may include the name of the film and television and the specific episode number, etc.; if the second target keyword is recognized, the video file processing system will continue to use the second keyword to obtain the second media data corresponding to the video file from multiple media repositories, and then use the multi-source data fusion technology to fuse all the second media data obtained by the query, obtain multiple fused media data, and further combine the multiple fused media data to obtain a fused media data set.
[0055] It should be understood that the target keyword here includes the second target keyword, and it can also be understood that the target keyword is the second target keyword, and only a distinction is made here.
[0056] In one embodiment, as Figure 4 shown, that is, in step S50, that is, in the step of sequentially performing weight matching calculation on each fused media data in the fused media data set to obtain the optimal fused media data corresponding to the video file, the following steps are included: S51. Sequentially perform weight matching calculation on each fused media data in the fused media data set and the metadata to obtain a comprehensive matching score corresponding to each fused media data; S52. Sort the fused media data in the fused media data set based on the comprehensive matching score to obtain a sorting result; S53. Select the fused media data with the highest comprehensive matching score from the sorting result as the optimal fused media data corresponding to the video file.
[0057] As an example, for each piece of fused media data in the fused media data set, perform weight matching calculation with the metadata of the video file extracted from the target storage device. It can be to perform weight matching calculation on each data item in the fused media data and the corresponding data item in the metadata, so as to obtain a comprehensive matching score corresponding to each fused media data, and then sort all the fused media data in the fused media data set based on the calculated comprehensive matching score to obtain a sorting result. Finally, select the fused media data with the highest comprehensive matching score from the sorting result and determine it as the optimal fused media data corresponding to the video file, so as to accurately screen out the film and television information most suitable for the video file from many fused media data, and improve the accuracy of film and television information matching.
[0058] In one embodiment, as Figure 5As shown, that is, in step S51, that is, in the process of sequentially performing weighted matching calculations on each piece of fused media asset data in the fused media asset dataset with the metadata to obtain the comprehensive matching score corresponding to each piece of fused media asset data, the following steps are included: S511. According to the preset priority of the metadata, divide the metadata into first-level metadata and second-level metadata, and set corresponding matching weights for the first-level metadata and the second-level metadata respectively; S512. For each piece of fused media asset data in the fused media asset dataset, perform the following steps: S512A. Calculate the first matching degree score between the fused media asset data and the first-level metadata; S512B. Calculate the second matching degree score between the fused media asset data and the second-level metadata; S512C. Obtain the first weighted score according to the first matching degree score and the matching weight corresponding to the first-level metadata; S512D. Obtain the second weighted score according to the second matching degree score and the matching weight corresponding to the second-level metadata; S512E. Obtain the comprehensive matching score of the fused media asset data according to the first weighted score and the second weighted score.
[0059] In this embodiment, the first-level metadata may include, but is not limited to, the release year and type (such as TV series, movies, or animations); the second-level metadata may include, but is not limited to, the director, actors, country, clarity, and audio-video type.
[0060] Among them, the priority of the first-level metadata is greater than that of the second-level metadata, and the matching weight corresponding to the first-level metadata is also greater than that of the second-level metadata. For example, the matching weight corresponding to the first-level metadata can be set to 0.7, and the matching weight corresponding to the second-level metadata can be set to 0.3. It should be understood that the priority of the above metadata and the matching weights corresponding to each level of metadata can be set as needed and are not limited here.
[0061] As an example, for each piece of fused media asset data in the fused media asset dataset, the comprehensive matching score is calculated according to the following steps: First, calculate the first matching degree score between the data items in the fused media asset data that match the first-level metadata, and the second matching degree score between the data items in the fused media asset data that match the second-level metadata. Then, multiply the first matching degree score of the first-level metadata by its corresponding matching weight to obtain the first weighted score; similarly, multiply the second matching degree score of the second-level metadata by its corresponding matching weight to obtain the second weighted score. Subsequently, add the first weighted score and the second weighted score to obtain the comprehensive matching score of the fused media asset data.
[0062] It should be noted that the above calculation process is only an example. In this application, the comprehensive matching score corresponding to each fused media asset data can also be obtained through other methods. For example, the matching degree score is added to the corresponding matching weight to obtain the corresponding weighted score, and further, the total matching score is obtained based on the weighted score, which is not limited here.
[0063] It should be understood that by dividing the matching calculation of the fused media asset data into two levels of first-level metadata and second-level metadata, the precise control of the priority of the media asset data is achieved. And this hierarchical matching strategy can maximize the influence of the core data, while taking into account other important attributes, and avoid the overall result from being inaccurate due to the mismatch of some secondary dimensions.
[0064] For example, assume that the metadata extracted from a video file includes the release year 2022, type movie, director A, starring B and C, and resolution 1080P. The fused media asset dataset contains three pieces of data. Calculate the matching degree scores based on the first-level metadata (such as release year and type, etc.) and the second-level metadata (such as director, starring, and resolution, etc.), and assign weights. The weight of the first-level metadata is 0.7, and the weight of the second-level metadata is 0.3.
[0065] For the first piece of fused media asset data, which includes a release year of 2021, type movie, director A, starring B and C, and resolution 1080P. In the calculation process of the first matching degree score, the type matching degree is 1 (fully matched). Since the release year is close to that of the video file, its matching degree can be calculated as 0.9 (assuming that close years are scored according to a certain ratio). Then the first matching degree score with the first-level metadata is (1 + 0.9) / 2 = 0.95, and the weighted score is 0.95 × 0.7 = 0.665. In the calculation process of the second matching degree score, the director, starring, and resolution are all fully matched, and their comprehensive matching degree is 1. The weighted score is 1 × 0.3 = 0.3. Finally, the comprehensive matching score corresponding to the first piece of fused media asset data is 0.665 + 0.3 = 0.965.
[0066] For the second piece of fused media asset data, which includes a release year of 2018, type movie, director A, starring B and C, and resolution 720P, the calculated comprehensive matching score corresponding to the second piece of fused media asset data is 0.425; and for the third piece of fused media asset data, which includes a release year of 2022, type movie, director A, starring B and C, and resolution 1080P, the calculated comprehensive matching score corresponding to the third piece of fused media asset data is 1. The calculation process here can refer to the calculation process of the first piece of fused media asset data. To avoid repetition, it will not be elaborated here.
[0067] It should be noted that the above is only an example and does not constitute a limitation to the present invention.
[0068] In one embodiment, as Figure 2 shown, the video file information includes the file name and / or path information corresponding to the video file, and the method further includes the following steps: S60. When the recognition result is that the target keyword is not recognized, use a target algorithm to extract a target extraction keyword from the file name and / or path information; S70. Query a fused media asset dataset from the multiple media asset libraries according to the target extraction keyword.
[0069] In this embodiment, the target algorithm may include, but is not limited to, a segmentation algorithm or other algorithms, etc.; the path information may include, but is not limited to, a path structure and other information, etc.
[0070] As an example, when the target recognition model fails to recognize the target keyword, it means that the pre-trained target recognition model cannot understand the naming rule and path information of the current video file, and thus cannot accurately recognize the target keyword or makes a wrong recognition. At this time, in order to avoid the situation that the film and television information of the video file cannot be accurately displayed due to the failure to recognize the target keyword, a pre-set target algorithm will be used to segment the file name and / or path information included in the video file information to obtain the target extraction keyword.
[0071] For example, for the video file name "Film_ABC_2023.mp4", using the segmentation algorithm according to rules such as character characteristics and semantic separation, it is segmented into several parts such as "Film", "ABC", "2023", etc., and used as the target extraction keyword.
[0072] Further, querying a fused media asset dataset from multiple media asset libraries according to the target extraction keyword, the specific process can refer to the description of the embodiment of step S30, and will not be repeated here to avoid redundancy.
[0073] In addition, the target recognition model will also record the file name and path structure at this time, and perform dynamic learning based on the current naming rule and path structure, so that when the corresponding file name and / or path information is recognized next time, the target extraction keyword can be accurately recognized, improving the fault tolerance of the video file processing system.
[0074] In one embodiment, as Figure 1 and 2 shown, that is, after step S50, that is, after storing the optimal fused media asset data corresponding to the video file in the database, it includes: S80. Display the optimal fusion media asset data corresponding to the video file in a target form on the user interface according to the category corresponding to the video file.
[0075] As an example, after storing the optimal fusion media asset data corresponding to the video file in the local database, the category corresponding to the video file can be further obtained, and the corresponding optimal fusion media asset data can be displayed in a target form on the user interface according to the category. For example, it can be displayed in the form of a poster wall on the user interface, so that users can intuitively and conveniently obtain the detailed film and television information of the video file, improving the user experience and the practicality of the film and television resource management system.
[0076] It should be understood that the user interface can be the interface on the display device, such as the display interface of various televisions or the display interface of a monitor, which is not limited here.
[0077] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0078] In a second aspect, a film and television information matching device for a video file is provided. The film and television information matching device for the video file corresponds one-to-one with the film and television information matching method for the video file in the above embodiments. As Figure 6 shown, the film and television information matching device for the video file includes a first acquisition module 101, an input module 102, a second acquisition module 103, an extraction module 104, and a storage module 105. The detailed descriptions of each functional module are as follows: The first acquisition module 101 is configured to acquire the video file information of the video file in the target storage device; The input module 102 is configured to input the video file information into a pre-trained target recognition model to obtain the recognition result output by the target recognition model; The second acquisition module 103 is configured to, when the recognition result is that a target keyword is recognized, query a fusion media asset data set from multiple media asset libraries according to the target keyword; The extraction module 104 is configured to extract the metadata corresponding to the video file from the target storage device; The storage module 105 is configured to perform a weight matching calculation on each fusion media asset data in the fusion media asset data set with the metadata in sequence to obtain the optimal fusion media asset data corresponding to the video file, and store the fusion media asset data corresponding to the video file in the database.
[0079] In one embodiment, the pre-trained target recognition model includes a pre-trained video recognition model and a video collection recognition model; the input module 102 is specifically configured to: Determine the category corresponding to the video file according to the video file information; When the category is the first category, input the video file information into the video recognition model to obtain a first recognition result output by the video recognition model; When the category is the second category, input the video file information into the video collection recognition model to obtain a second recognition result output by the video collection recognition model.
[0080] In one embodiment, the target keyword includes a first target keyword; the second acquisition module 103 is specifically configured to: When the recognition result is the first recognition result, determine whether the first recognition result recognizes the first target keyword; If the first target keyword is recognized, use the first target keyword to query all first media data from multiple media libraries respectively; Use multi-source data fusion technology to fuse all the first media data to obtain the fused media data set.
[0081] In one embodiment, the target keyword further includes a second target keyword; the second acquisition module 103 is specifically further configured to: When the recognition result is the second recognition result, determine whether the second recognition result recognizes the second target keyword; If the second target keyword is recognized, use the second target keyword to query all second media data from multiple media libraries respectively; Use multi-source data fusion technology to fuse all the second media data to obtain the fused media data set.
[0082] In one embodiment, the storage module 105 is specifically configured to: Perform weight matching calculation on each fused media data in the fused media data set with the metadata in sequence to obtain a comprehensive matching score corresponding to each fused media data; Sort the fused media data in the fused media data set based on the comprehensive matching score to obtain a sorting result; Select the fused media data with the highest comprehensive matching score from the sorting result as the optimal fused media data corresponding to the video file.
[0083] In one embodiment, the storage module 105 is specifically further configured to: Divide the metadata into primary metadata and secondary metadata according to the preset priority of the metadata, and set corresponding matching weights for the primary metadata and the secondary metadata respectively; For each piece of integrated media asset data in the integrated media asset dataset, perform the following steps: Calculate a first matching degree score between the integrated media asset data and the primary metadata; Calculate a second matching degree score between the integrated media asset data and the secondary metadata; Obtain a first weighted score according to the first matching degree score and the matching weight corresponding to the primary metadata; Obtain a second weighted score according to the second matching degree score and the matching weight corresponding to the secondary metadata; Obtain a comprehensive matching score of the integrated media asset data according to the first weighted score and the second weighted score.
[0084] In an embodiment, the video file information includes the file name and / or path information corresponding to the video file; the second acquisition module 103 is further configured to: When the recognition result is that the target keyword is not recognized, extract a target extraction keyword from the file name and / or path information by using a target algorithm; Query the integrated media asset dataset from the multiple media asset libraries according to the target extraction keyword.
[0085] For the specific definition of the film and television information matching device for video files, reference can be made to the definition of the film and television information matching method for video files in the above text, which will not be elaborated here. Each module in the above film and television information matching device for video files can be implemented in whole or in part by software, hardware and their combination. The above modules can be embedded in the processor of the smart terminal in hardware form or be independent of it, or can be stored in the memory of the smart terminal in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0086] In a third aspect, a smart terminal is provided, and its internal structure diagram can be as Figure 7As shown in the figure. The intelligent terminal includes a processor, a memory, a network interface, and a database connected via a system bus. Among them, the processor of the intelligent terminal is used to provide computing and control capabilities. The memory of the intelligent terminal includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the intelligent terminal is used to store all data during the process of the video information matching method for running video files. The network interface of the intelligent terminal is used to connect and communicate with external terminals. When the computer program is executed by the processor, it implements the video information matching method for various video files described in the first aspect above.
[0087] In one embodiment, an intelligent terminal is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented: Obtain the video file information of the video file in the target storage device; Input the video file information into a pre-trained target recognition model to obtain the recognition result output by the target recognition model; When the recognition result is that a target keyword is recognized, query a fusion media asset data set from multiple media asset libraries according to the target keyword; Extract the metadata corresponding to the video file from the target storage device; Perform weighted matching calculations on each fusion media asset data in the fusion media asset data set with the metadata in turn to obtain the optimal fusion media asset data corresponding to the video file, and store the optimal fusion media asset data in the database.
[0088] In one embodiment, a computer-readable storage medium is provided. A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, the following steps are implemented: Obtain the video file information of the video file in the target storage device; Input the video file information into a pre-trained target recognition model to obtain the recognition result output by the target recognition model; When the recognition result is that a target keyword is recognized, query a fusion media asset data set from multiple media asset libraries according to the target keyword; Extract the metadata corresponding to the video file from the target storage device; Perform weighted matching calculations on each fusion media asset data in the fusion media asset data set with the metadata in turn to obtain the optimal fusion media asset data corresponding to the video file, and store the optimal fusion media asset data in the database.
[0089] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. This computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0090] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.
[0091] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A method for matching film and television information of a video file, characterized in that: include: Obtain video file information of video files in a target storage device; Inputting the video file information into a pre-trained target recognition model to obtain a recognition result output by the target recognition model; When the recognition result is that the target keyword is recognized, querying from multiple media resource libraries to obtain a fused media resource data set according to the target keyword; Extracting metadata corresponding to the video file from the target storage device; Each fused media asset data in the fused media asset data set is sequentially matched with the metadata to obtain the optimal fused media asset data corresponding to the video file, and the optimal fused media asset data is stored in a database.
2. The video information matching method of a video file according to claim 1, characterized in that: The pre-trained object recognition model includes a pre-trained video recognition model and a video collection recognition model; The step of inputting the video file information into a pre-trained target recognition model to obtain a recognition result output by the target recognition model includes: Determining a category corresponding to the video file according to the video file information; When the category is the first category, inputting the video file information into the video recognition model to obtain a first recognition result output by the video recognition model; When the category is the second category, the video file information is input into the video collection recognition model to obtain a second recognition result output by the video collection recognition model.
3. The video information matching method of a video file according to claim 2, characterized in that: The target keywords include a first target keyword; When the recognition result is that the target keyword is recognized, querying from multiple media resource libraries according to the target keyword to obtain a fused media resource data set includes: When the recognition result is the first recognition result, determining whether the first recognition result recognizes the first target keyword; If the first target keyword is identified, all the first media asset data are queried from the plurality of media asset libraries using the first target keyword; All the first media asset data are fused using multi-source data fusion technology to obtain the fused media asset data set.
4. The video information matching method of a video file according to claim 2, characterized in that: The target keyword includes a second target keyword; When the recognition result is that the target keyword is recognized, querying from multiple media resource libraries according to the target keyword to obtain a fused media resource data set includes: When the recognition result is the second recognition result, determining whether the second recognition result recognizes the second target keyword; If the second target keyword is identified, all the second media asset data are queried from the plurality of media asset libraries using the second target keyword; All the second media asset data are fused using multi-source data fusion technology to obtain the fused media asset data set.
5. The video information matching method of a video file according to claim 1, characterized in that: The step of performing weight matching calculation on each fused media asset data in the fused media asset data set and the metadata in sequence to obtain the optimal fused media asset data corresponding to the video file includes: Performing weighted matching calculation on each fused media asset data in the fused media asset data set and the metadata in turn to obtain a comprehensive matching score corresponding to each fused media asset data; Based on the comprehensive matching score, sorting the fused media asset data in the fused media asset data set to obtain a sorting result; The fused media asset data with the highest comprehensive matching score is selected from the sorting results as the optimal fused media asset data corresponding to the video file.
6. The video information matching method of a video file according to claim 5, characterized in that: The step of performing weighted matching calculation on each fused media asset data in the fused media asset data set and the metadata in sequence to obtain a comprehensive matching score corresponding to each fused media asset data includes: According to the priority of the pre-set metadata, the metadata is divided into primary metadata and secondary metadata, and corresponding matching weights are set for the primary metadata and the secondary metadata respectively; For each fused media asset data in the fused media asset data set, the following steps are performed: Calculating a first matching score between the fused media asset data and the primary metadata; Calculating a second matching degree score between the fused media asset data and the secondary metadata; Obtaining a first weighted score according to the first matching degree score and a matching weight corresponding to the primary metadata; Obtaining a second weighted score according to the second matching degree score and a matching weight corresponding to the secondary metadata; A comprehensive matching score of the fused media asset data is obtained according to the first weighted score and the second weighted score.
7. The video information matching method of a video file according to any one of claims 1 to 6, characterized in that: The video file information includes the file name and / or path information corresponding to the video file, and the method further includes: When the recognition result is that the target keyword is not recognized, extracting the target extraction keyword from the file name and / or path information using a target algorithm; Keywords are extracted according to the target, and a fused media asset data set is obtained by querying from the multiple media asset libraries.
8. A video information matching device for a video file, characterized in that: include: A first acquisition module, used to acquire video file information of video files in a target storage device; An input module, used to input the video file information into a pre-trained target recognition model to obtain a recognition result output by the target recognition model; A second acquisition module is used for, when the recognition result is that a target keyword is recognized, querying from multiple media resource libraries to obtain a fused media resource data set according to the target keyword; An extraction module, used for extracting metadata corresponding to the video file from the target storage device; The storage module is used to perform weight matching calculation on each fused media asset data in the fused media asset data set and the metadata in turn to obtain the optimal fused media asset data corresponding to the video file, and store the fused media asset data corresponding to the video file in a database.
9. An intelligent terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method for matching film and television information of a video file according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for matching film and television information of a video file according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Video file information association method and device, equipment and storage medium
CN121012957A
Data fusion method and device, storage medium and electronic equipment
CN121542991A