Picture and video matching method and system based on multi-modal fusion

By using a multimodal fusion image and video matching system, and combining social network interest sets and frame matching features for customized adjustments, the problem of not being able to meet users' personalized needs in existing technologies has been solved, and more efficient and accurate matching results have been achieved.

CN120997543AInactive Publication Date: 2025-11-21HANGZHOU YINSHAN TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511517798.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2025-11-21
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing image and video matching technologies struggle to fully uncover the deep semantic relationships in cross-modal data, lack the ability to dynamically perceive users' personalized needs, and are unable to generate targeted video adjustment strategies. Furthermore, they struggle to parse the semantic logic of multi-keyword combinations when users submit unstructured update requests, impacting the efficiency and accuracy of matching optimization.

Method used

By using a multimodal fusion-based image and video matching system, matching video sets are generated using matching tag association data. Social network interest sets are combined to predict and recommend scenarios. Frame matching features and key point response videos are used to customize and adjust the videos, guiding users to preview the customized video sets. At the same time, keywords are decomposed and combined in the matching video sets to create updated combination models in response to user update requests.

Benefits of technology

It improves the accuracy of matching results and user experience, enhances the applicability and suitability of the system, improves matching accuracy in complex virtual scenarios, meets users' personalized needs, and optimizes matching results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997543A_ABST
    Figure CN120997543A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data matching, in particular to a picture and video matching method and system based on multi-modal fusion, and the method comprises the steps: when a matching demander recognizes a matching tag on a to-be-matched picture, obtaining a matched video set according to the to-be-matched picture matching effect data associated with the matching tag; according to the matching video set, guiding the matching demander to update the matching effect of the to-be-matched picture; wherein the matching effect recommendation scene of the to-be-matched picture needs to be predicted according to the social network interest set of the matching demander; further obtaining a recommended target scene; customizing and adjusting the matched video set according to the recommended target scene; and guiding the matching demander to preview the customized and adjusted matching video set. According to the method, the reality sense of the scene is stronger, the phenomenon of scene distortion occurs as much as possible, and the matching accuracy can be improved in a huge and complex virtual scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data matching technology, and in particular to a method and system for image and video matching based on multimodal fusion. Background Technology

[0002] Existing image and video matching technologies typically rely on single-modal features (such as visual or text labels) for association analysis, making it difficult to fully explore the deep semantic relationships across modal data. Especially when users actively initiate matching requests, the system can only provide basic video retrieval results, lacking the ability to dynamically perceive users' personalized needs. This results in insufficient alignment between the matching results and the user's actual application scenario, failing to effectively support subsequent interactive optimization operations.

[0003] Traditional solutions typically employ static recommendation mechanisms when guiding users to update image matching results, failing to provide contextualized guidance based on user behavior and preferences. For example, the system does not consider the correlation between users' social interests and target scenes, making it unable to generate targeted video adjustment strategies. Furthermore, existing methods have limited dynamic reconstructive capabilities for video content, making it difficult to achieve adaptive scene reconstruction through keyframe analysis, thus restricting the accuracy and operability of recommended target scenes.

[0004] Deep response mechanisms for complex user commands also have limitations. When users submit unstructured update requests, existing technologies struggle to effectively parse the semantic logic of multi-keyword combinations and lack the ability to model multi-scale relationships between matching parameters. This prevents the system from generating a visually guided model that integrates parameter relationships, making it difficult for users to intuitively perceive the global impact of update operations, ultimately affecting the efficiency and accuracy of matching optimization. Summary of the Invention

[0005] To achieve the above objectives, this application provides the following technical solution: According to a first aspect of the present invention, the present invention claims protection for an image and video matching system based on multimodal fusion, comprising: The matching video unit is used to obtain a set of matching videos based on the matching effect data of the images to be matched associated with the matching tags when the matching personnel identify the matching tags on the images to be matched. The update unit is used to guide the matching personnel to update the matching effect of the images to be matched based on the set of matching videos. The update unit, based on the set of matching videos, guides users to update the matching results for the images to be matched, including: Based on the social network interest sets of people with matching needs, predict the matching effect of the image to be matched and recommend scenarios; Based on the matching effect of the image to be matched, a scene is recommended to obtain the recommended target scene; Based on the recommended target scenario, the matching video set is customized and adjusted; Guide users with matching needs to preview the customized and adjusted set of matching videos.

[0006] Furthermore, the updating unit recommends scenes based on the matching effect of the image to be matched, and obtains the recommended target scene, including: The matching effect of the images to be matched is recommended in the scene by video frame representation, and multiple video frames are obtained; Frame graph analysis is performed on the correlation between the metadata of each video frame and between pairs of video frames to obtain frame graph matching features; Based on frame-matching features, key point responses are generated in the video, and multiple corresponding key points are generated in similar scenes. Based on the frame matching features, video frames are associated with each key point. Similar scenes obtained by associating each video frame with each key point are used as recommended target scenes.

[0007] Furthermore, the updating unit customizes and adjusts the matching video set according to the recommended target scene, including: List the key points in the recommended target scene in descending order of the total number of associated video frames to obtain the key point set; Multiple partition subsets are determined from the set of key points; wherein the difference in the total number of video frames associated with adjacent key points in the same partition subset does not exceed the difference threshold or they belong to associated key points in the recommended target scene; Each partition subset is polled in turn. During each poll, the model is customized based on the video frames associated with each key point in the polled partition subset, and then associated with the polled partition subset. After polling each partition subset, the model customizations obtained in each polling are listed according to the order of the associated partition subsets in the keypoint set to obtain the customization set; The customized adjustment video is obtained; wherein, the customized adjustment video includes: matching the video set and displaying the customized models in the customized set one by one in the set order, and the duration of each display is the duration value associated with the order of the displayed model customization in the customized set; Based on the customized video, the matching video set is customized and adjusted.

[0008] Furthermore, after the updating unit guides the matching personnel to update the matching effect of the image to be matched based on the matching video set, it also includes: The deep update unit is used for: When the target content of the update message associated with the matching requester's input in the matching video set is blank, the update message is decomposed into keywords to obtain multiple matching instructions; The matching instructions are combined to obtain multiple keyword combinations; among them, the type distribution of each matching instruction in the same keyword combination matches the standard type distribution, and the minimum number of matching parameters in the matching parameter association table of each matching instruction in the same keyword combination exceeds the threshold of the number of matching parameters. Each keyword combination is polled in turn. During each poll, the video is obtained based on the distribution of matching instructions in the polled keyword combinations. Matching parameter query conditions are obtained based on the matching instructions in the polled keyword combinations. The target matching parameters are then retrieved from the matching parameter library based on the matching parameter query conditions. After polling each keyword combination, the target matching parameters retrieved during each polling are combined to obtain the combined matching parameters; Based on the combined matching parameters, respond with an update message.

[0009] Furthermore, the deep update unit performs matching parameter combination on the target matching parameters queried during each polling, including: Based on the multi-scale relationships between the pairwise target matching parameters retrieved during each polling, an updated combined model and an updated guiding video are created; The target matching parameters retrieved during each polling are associated with the updated combination model; when the matching personnel are guided by the update guidance video to update the matching effect based on the updated combination model after all target matching parameters are associated, the matching personnel can perceive all multi-scale relationships. The updated combined model, which associates all target matching parameters, and the updated guidance video are used as target matching parameters.

[0010] According to a second aspect of the present invention, the present invention claims protection for an image and video matching method based on multimodal fusion, comprising: When the matching personnel identify the matching tags on the images to be matched, they obtain a set of matching videos based on the matching effect data of the images to be matched associated with the matching tags. Based on the set of matching videos, guide the personnel who need to match to update the matching effect of the images to be matched; The step of guiding users to update the matching results for images based on the matching video set includes: Based on the social network interest sets of people with matching needs, predict the matching effect of the image to be matched and recommend scenarios; Based on the matching effect of the image to be matched, a scene is recommended to obtain the recommended target scene; Based on the recommended target scenario, the matching video set is customized and adjusted; Guide users with matching needs to preview the customized and adjusted set of matching videos.

[0011] Furthermore, the step of recommending scenarios based on the matching effect of the image to be matched, and obtaining the recommended target scenario, includes: The matching effect of the images to be matched is recommended in the scene by video frame representation, and multiple video frames are obtained; Frame graph analysis is performed on the correlation between the metadata of each video frame and between pairs of video frames to obtain frame graph matching features; Based on frame-matching features, key point responses are generated in the video, and multiple corresponding key points are generated in similar scenes. Based on the frame matching features, video frames are associated with each key point. Similar scenes obtained by associating each video frame with each key point are used as recommended target scenes.

[0012] Furthermore, the process of customizing and adjusting the matching video set based on the recommended target scene includes: List the key points in the recommended target scene in descending order of the total number of associated video frames to obtain the key point set; Multiple partition subsets are determined from the set of key points; wherein the difference in the total number of video frames associated with adjacent key points in the same partition subset does not exceed the difference threshold or they belong to associated key points in the recommended target scene; Each partition subset is polled in turn. During each poll, the model is customized based on the video frames associated with each key point in the polled partition subset, and then associated with the polled partition subset. After polling each partition subset, the model customizations obtained in each polling are listed according to the order of the associated partition subsets in the keypoint set to obtain the customization set; The customized adjustment video is obtained; wherein, the customized adjustment video includes: matching the video set and displaying the customized models in the customized set one by one in the set order, and the duration of each display is the duration value associated with the order of the displayed model customization in the customized set; Based on the customized video, the matching video set is customized and adjusted.

[0013] Furthermore, after guiding the matching personnel to update the matching effect of the image to be matched based on the matching video set, the process also includes: When the target content of the update message associated with the matching requester's input in the matching video set is blank, the update message is decomposed into keywords to obtain multiple matching instructions; The matching instructions are combined to obtain multiple keyword combinations; among them, the type distribution of each matching instruction in the same keyword combination matches the standard type distribution, and the minimum number of matching parameters in the matching parameter association table of each matching instruction in the same keyword combination exceeds the threshold of the number of matching parameters. Each keyword combination is polled in turn. During each poll, the video is obtained based on the distribution of matching instructions in the polled keyword combinations. Matching parameter query conditions are obtained based on the matching instructions in the polled keyword combinations. The target matching parameters are then retrieved from the matching parameter library based on the matching parameter query conditions. After polling each keyword combination, the target matching parameters retrieved during each polling are combined to obtain the combined matching parameters; Based on the combined matching parameters, respond with an update message.

[0014] Furthermore, the step of combining the target matching parameters retrieved during each polling includes: Based on the multi-scale relationships between the pairwise target matching parameters retrieved during each polling, an updated combined model and an updated guiding video are created; The target matching parameters retrieved during each polling are associated with the updated combination model; when the matching personnel are guided by the update guidance video to update the matching effect based on the updated combination model after all target matching parameters are associated, the matching personnel can perceive all multi-scale relationships. The updated combined model, which associates all target matching parameters, and the updated guidance video are used as target matching parameters.

[0015] This application relates to the field of data matching technology, and in particular to a method and system for image and video matching based on multimodal fusion. When a user identifies matching tags on an image to be matched, a set of matching videos is obtained based on the matching effect data of the image to be matched associated with the matching tags. Based on the set of matching videos, the user is guided to update the matching effect of the image to be matched. This involves predicting recommended scenarios for the matching effect of the image to be matched based on the user's social network interest set, and further obtaining recommended target scenarios. The set of matching videos is then customized and adjusted based on the recommended target scenarios. The user is then guided to preview the customized and adjusted set of matching videos. This invention enhances the realism of the scene, minimizing scene distortion, and can improve matching accuracy in large and complex virtual scenes. Attached Figure Description

[0016] Figure 1 A structural block diagram of an image and video matching system based on multimodal fusion, as claimed in an embodiment of this application; Figure 2This is a flowchart illustrating the process of an image and video matching method based on multimodal fusion, as claimed in an embodiment of this application. Detailed Implementation

[0017] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0018] The terms "first," "second," and "third" in this application are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first," "second," or "third" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified. All directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of this application are only used to explain the relative positional relationships and movement of components in a specific posture (as shown in the figures). If the specific posture changes, the directional indication also changes accordingly. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.

[0019] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0020] This invention provides an image and video matching system based on multimodal fusion, such as... Figure 1 As shown, it includes: The matching video unit 1 is used to obtain a set of matching videos based on the matching effect data of the image to be matched associated with the matching tags when the matching personnel identify the matching tags on the image to be matched. Update Unit 2 is used to guide the matching personnel to update the matching effect of the images to be matched based on the set of matching videos; The update unit, based on the set of matching videos, guides users to update the matching results for the images to be matched, including: Based on the social network interest sets of people with matching needs, predict the matching effect of the image to be matched and recommend scenarios; Based on the matching effect of the image to be matched, a scene is recommended to obtain the recommended target scene; Based on the recommended target scenario, the matching video set is customized and adjusted; Guide users with matching needs to preview the customized and adjusted set of matching videos.

[0021] In the above technical solution, the matching demanders are consumers who purchase the images to be matched; each image to be matched is assigned a matching tag when it is taken, and the matching tag is associated with the matching effect data of the image to be matched. The matching effect data of the image to be matched includes at least: shooting location, shooting log, detection record, responsible personnel data, shooting date, etc.

[0022] The matching video set displays the matching effect data of the images to be matched to the users who need matching; therefore, the matching video set can guide users to update the matching effect of the images to be matched.

[0023] The social network interest set of the matching applicants should include at least: their age and their image and video matching history based on multimodal fusion; the recommended scenarios for the matching effect of the image to be matched should include at least: the game progress stage and game progress conditions that the matching applicants may recommend to affect the matching effect of the image to be matched; the social network interest set will reflect the recommended scenarios for the matching effect of the image to be matched, therefore, the latter can be predicted based on the former. For example, if the image and video matching history based on multimodal fusion reflects the relevant data of the game background scene stage that the matching applicants have historically frequently used to match the image to be matched, then the predicted recommended scenarios for the matching effect of the image to be matched are the game progress stage that the matching applicants may recommend to affect the matching effect of the image to be matched.

[0024] Based on the matching effect of the image to be matched, a recommended scene is obtained, which is the feature representation of the recommended scene based on the matching effect of the image to be matched.

[0025] The recommended target scenarios reflect the recommended scenarios for the matching effect of the images to be matched by the users. Therefore, the matching video set can be customized and adjusted based on the recommended target scenarios. For example, if the recommended target scenarios reflect the game progress stages that the users may recommend that affect the matching effect of the images to be matched, such as the game background scene stage, then the matching video set can be customized to first display the relevant data of the images to be matched in the game background scene stage, so that the users can preview it first.

[0026] After the matching video set has been customized and adjusted, guide the users who need matching to preview the customized and adjusted matching video set.

[0027] This application obtains a set of matching videos based on the matching effect data of the images to be matched associated with matching tags. Based on the set of matching videos, it guides users to update the matching effect of the images to be matched. Specifically, during the guidance, it predicts the recommended scenarios for the matching effect of the images to be matched based on the social network interest set of the users and obtains the recommended target scenarios. Based on the recommended target scenarios, it customizes and adjusts the set of matching videos and guides users to preview the customized and adjusted set of matching videos. Under the guidance, users can intuitively, comprehensively, and efficiently preview the data and perform image and video matching based on multimodal fusion.

[0028] In one embodiment, the updating unit recommends a scene based on the matching effect of the image to be matched, and obtains a recommended target scene, including: The matching effect of the images to be matched is recommended in the scene by video frame representation, and multiple video frames are obtained; Frame graph analysis is performed on the correlation between the metadata of each video frame and between pairs of video frames to obtain frame graph matching features; Based on frame-matching features, key point responses are generated in the video, and multiple corresponding key points are generated in similar scenes. Based on the frame matching features, video frames are associated with each key point. Similar scenes obtained by associating each video frame with each key point are used as recommended target scenes.

[0029] In the above technical solution, when representing video frames, multiple data features of the recommended scene of the matching effect of the image to be matched are extracted, and each data feature is classified into multiple feature sets of classification categories. Each feature set is represented by cloud to obtain video frames. Among them, the classification categories include at least: belonging to time features, belonging to location features, and belonging to features related to the stage of the game background scene.

[0030] Video frames have metadata, which is represented by the feature set of the video frame as a classification category. The correlation between the metadata of any two video frames includes at least the following: they can reflect the same recommendation type (e.g., if the classification categories are both from recommended game manufacturer data and both from recommended game processing data, then the correlation is that they reflect the same recommendation type as games), similar trigger times, and connection relationships between game progress stages. When performing frame graph analysis on video frames and correlations, the metadata of the video frame, the number of features in the video frame, and the correlation are represented in vector form to obtain frame graph matching features.

[0031] There are multiple key points in similar scenarios, each key point representing a recommendation point; the frame-map matching feature is associated with pre-set key point response videos. The key point response videos are the key points associated with the recommendation points of the matching users reflected by the frame-map matching feature, so that they are waiting for data association. For example, if the metadata of the video frames used to construct the frame-map matching feature is that they belong to the same recommended game manufacturer and the relevance is that they reflect the same recommendation type as games, then the recommendation point of the matching users is the game manufacturer of the image to be matched, and the key point response video is the key point that responds to the game manufacturer.

[0032] The frame matching feature is also associated with pre-set video frame association videos. The video frame association video indicates how to associate the video frame with key points, including: associating the video frame with the key points represented by the features in the video frame that match the recommended points of the matching needs of the personnel. For example, if the data that constitutes the frame matching feature is that the metadata of a certain video frame belongs to the same recommended game manufacturer, then the video frame association video is the key point that associates the video frame with the game manufacturer.

[0033] After each video frame is associated with a key point, similar scenes after the association of each video frame with each key point are used as the recommended target scenes. In this way, the more video frames associated with the same key point, the higher the degree to which the recommended point represented by the key point matches the recommendation of the user.

[0034] In this embodiment of the invention, when obtaining the recommended target scene, video frames and similar scenes are introduced. Based on the key point response video associated with the frame image matching features, the video frames are associated with the video, and the corresponding key points in the similar scene are responded to. The corresponding video frames are then associated with each key point. The similar scene after each video frame is associated with each key point is used as the recommended target scene. This allows the recommended target scene to fully reflect the recommended scene for the matching effect of the matching needs of the matching personnel, which greatly improves the suitability and accuracy of the recommended target scene and enhances the applicability of the system.

[0035] In one embodiment, the updating unit customizes and adjusts the matching video set according to the recommended target scene, including: List the key points in the recommended target scene in descending order of the total number of associated video frames to obtain the key point set; Multiple partition subsets are determined from the set of key points; wherein the difference in the total number of video frames associated with adjacent key points in the same partition subset does not exceed the difference threshold or they belong to associated key points in the recommended target scene; Each partition subset is polled in turn. During each poll, the model is customized based on the video frames associated with each key point in the polled partition subset, and then associated with the polled partition subset. After polling each partition subset, the model customizations obtained in each polling are listed according to the order of the associated partition subsets in the keypoint set to obtain the customization set; The customized adjustment video is obtained; wherein, the customized adjustment video includes: matching the video set and displaying the customized models in the customized set one by one in the set order, and the duration of each display is the duration value associated with the order of the displayed model customization in the customized set; Based on the customized video, the matching video set is customized and adjusted.

[0036] In the above technical solution, the difference threshold can be 3. The more video frames associated with the same key point, the higher the degree to which the recommendation point represented by the key point matches the recommendation of the user. Therefore, ensuring that the difference in the total number of video frames associated with adjacent key points in the same subset does not exceed the difference threshold can make the recommendation points represented by adjacent key points in the same subset match the recommendation of the user to a similar degree, and the video frames on them are necessary to obtain the same model customization. Associated key points are pre-set on the recommendation target scene. When key points are associated key points, it means that the recommendation points represented by the key points belong to the same recommendation type, such as: the recommendation points represented by two key points are recommended game manufacturers and recommended game transportation data, which belong to recommended games. Therefore, ensuring that adjacent key points in the same subset are associated key points can also make the video frames on adjacent key points in the same subset necessary to obtain the same model customization.

[0037] When obtaining the model customization, ensure that all features in the video frames associated with each key point in the polled subset of partitions can be displayed within the same model customization; after obtaining the model customization, associate it with the polled subset of partitions.

[0038] When listing model customizations, for example, if the order of the model customization in the associated partition subset within the keypoint set is 2, then the order of the model customization in the customization set is also 2.

[0039] The duration value is negatively correlated with the order of model customization in the customization set. The smaller the order, the higher the position of the partition subset in the customization set, the higher the overall degree of recommendation by matching users, and the longer the customization retention time. In the customized adjustment video, the matching video set displays the model customization in the customization set one by one according to the set order. The retention time of each display is the duration value associated with the order of the displayed model customization in the customization set.

[0040] Based on the customized adjustment videos, the matching video set is customized and adjusted so that the adjusted matching video set is adapted to the recommendation scenario that reflects the matching needs of the users in the target recommendation scenario.

[0041] In this embodiment of the invention, when customizing and adjusting a set of matched videos, a set of key points is introduced. Multiple partition subsets are determined from the set of key points. Each partition subset is polled in turn. During each poll, the model customization is obtained based on the video frames associated with each key point in the polled partition subset. A customized set is determined, and a customized adjustment video is obtained based on the customized adjustment video. Based on the customized adjustment video, the set of matched videos is customized and adjusted, which greatly improves the rationality of customizing and adjusting the set of matched videos.

[0042] In one embodiment, after the updating unit guides the matching requester to update the matching effect of the image to be matched based on the matching video set, it further includes: The deep update unit is used for: When the target content of the update message associated with the matching requester's input in the matching video set is blank, the update message is decomposed into keywords to obtain multiple matching instructions; The matching instructions are combined to obtain multiple keyword combinations; among them, the type distribution of each matching instruction in the same keyword combination matches the standard type distribution, and the minimum number of matching parameters in the matching parameter association table of each matching instruction in the same keyword combination exceeds the threshold of the number of matching parameters. Each keyword combination is polled in turn. During each poll, the video is obtained based on the distribution of matching instructions in the polled keyword combinations. Matching parameter query conditions are obtained based on the matching instructions in the polled keyword combinations. The target matching parameters are then retrieved from the matching parameter library based on the matching parameter query conditions. After polling each keyword combination, the target matching parameters retrieved during each polling are combined to obtain the combined matching parameters; Based on the combined matching parameters, respond with an update message.

[0043] In the above technical solution, the matching personnel can also input update messages, which are the target content associated with the matching video set. If the target content is blank, it means that the update has failed and a deep update is required. First, the update message is broken down into multiple matching instructions, which include at least: message type, message content range, etc.

[0044] The type distribution of each matching instruction in the keyword combination indicates the types of matching instructions available. A pre-defined standard type distribution indicates the keyword types that can simultaneously serve as query conditions for matching parameters. For example, standard types include message content trigger time and message content trigger location. A pre-defined matching parameter association table records numerous data nodes recorded by the image manufacturer during the game progress of the image to be matched. Each data node represents a game progress event. The minimum number of matching parameters for each matching instruction in the matching parameter association table is the minimum value of the number of matching parameters for that instruction. The number of matching parameters for an instruction in the matching parameter association table is the number of data nodes related to that instruction. For example, if the matching instruction is "Message content trigger location is game library," then the related data nodes represent game progress videos within the game library. Ensuring that the minimum number of matching parameters for each matching instruction in the same keyword combination exceeds a matching parameter threshold allows query conditions obtained from the matching instructions in the keyword combination to be used to retrieve the target matching parameters.

[0045] The distribution of matching instructions in the keyword combination is associated with conditions for obtaining videos. These conditions indicate how to obtain the matching parameter query conditions for the target matching parameters based on the matching instructions in the keyword combination. For example, if the distribution of matching instructions is the content trigger time of the message and the content trigger location of the message, then the matching parameter query conditions are to retrieve the specific data of game progress events whose trigger time is the same as or similar to the content trigger time of the message and whose trigger location is the same as or similar to the content trigger location of the message. The matching parameter library is associated with the matching parameter association table, which records the specific data of the game progress events represented by different data nodes in the matching parameter association table.

[0046] After polling each keyword combination, the target matching parameters retrieved during each polling are combined to obtain the combined matching parameters.

[0047] Finally, based on the combined matching parameters, the system responds to update messages and can display the combined matching parameters to those who require matching within the set of matched videos, thus enabling the system to respond to update messages from those who require matching.

[0048] This application performs a deep update when the matching video set cannot meet the update messages of users with matching needs. It introduces keyword combinations, obtains matching parameter query conditions based on the keyword combinations, and retrieves the target matching parameters from the matching parameter library based on the matching parameter query conditions. After polling each keyword combination, it performs matching parameter combinations on the target matching parameters retrieved in each polling, and responds to update messages based on the combined matching parameters. This greatly improves the experience of users with matching needs and further enhances the applicability of the system.

[0049] In one embodiment, the deep update unit performs matching parameter combination on the target matching parameters queried during each polling, including: Based on the multi-scale relationships between the pairwise target matching parameters retrieved during each polling, an updated combined model and an updated guiding video are created; The target matching parameters retrieved during each polling are associated with the updated combination model; when the matching personnel are guided by the update guidance video to update the matching effect based on the updated combination model after all target matching parameters are associated, the matching personnel can perceive all multi-scale relationships. The updated combined model, which associates all target matching parameters, and the updated guidance video are used as target matching parameters.

[0050] In the above technical solution, the multi-scale relationship includes at least: the triggering time relationship and triggering location relationship of game progress events in the target matching parameters; based on the multi-scale relationship, an updated combination model and an updated guidance video are created; the updated combination model is a template model that can match the video display of the target matching parameters. For example, if the target matching parameter is a set of video interests in the game library, then the template model has an associated region for associating the set of video interests in the game library, and this associated region can be a three-dimensional model of the game library; the updated guidance video is a video that guides the matching requester to preview the updated combination model, which constrains the order in which the matching requester previews different contents in the updated combination model during the guidance. The target matching parameters retrieved during each polling are associated with the updated combination model. During association, the corresponding associated regions are found in the updated combination model based on the trigger time and location of the target matching parameters, and the association is performed. When the matching personnel are guided to update the matching effect based on the updated combination model after all target matching parameters are associated, according to the update guidance video, the matching personnel can perceive all multi-scale relationships. Finally, the updated combination model after all target matching parameters are associated and the update guidance video are used as target matching parameters. When using the target matching parameters, the matching personnel are guided to preview the updated combination model after all target matching parameters are associated, according to the update guidance video.

[0051] This invention provides an image and video matching method based on multimodal fusion, such as... Figure 2As shown, it includes: S1. When the matching personnel identify the matching tags on the images to be matched, they obtain a set of matching videos based on the matching effect data of the images to be matched associated with the matching tags. S2. Based on the set of matching videos, guide the personnel who need to match to update the matching effect of the images to be matched; The step of guiding users to update the matching results for images based on the matching video set includes: Based on the social network interest sets of people with matching needs, predict the matching effect of the image to be matched and recommend scenarios; Based on the matching effect of the image to be matched, a scene is recommended to obtain the recommended target scene; Based on the recommended target scenario, the matching video set is customized and adjusted; Guide users with matching needs to preview the customized and adjusted set of matching videos.

[0052] The step of recommending scenarios based on the matching effect of the image to be matched, and obtaining the recommended target scenario, includes: The matching effect of the images to be matched is recommended in the scene by video frame representation, and multiple video frames are obtained; Frame graph analysis is performed on the correlation between the metadata of each video frame and between pairs of video frames to obtain frame graph matching features; Based on frame-matching features, key point responses are generated in the video, and multiple corresponding key points are generated in similar scenes. Based on the frame matching features, video frames are associated with each key point. Similar scenes obtained by associating each video frame with each key point are used as recommended target scenes.

[0053] The process of customizing and adjusting the matching video set based on the recommended target scene includes: List the key points in the recommended target scene in descending order of the total number of associated video frames to obtain the key point set; Multiple partition subsets are determined from the set of key points; wherein the difference in the total number of video frames associated with adjacent key points in the same partition subset does not exceed the difference threshold or they belong to associated key points in the recommended target scene; Each partition subset is polled in turn. During each poll, the model is customized based on the video frames associated with each key point in the polled partition subset, and then associated with the polled partition subset. After polling each partition subset, the model customizations obtained in each polling are listed according to the order of the associated partition subsets in the keypoint set to obtain the customization set; The customized adjustment video is obtained; wherein, the customized adjustment video includes: matching the video set and displaying the customized models in the customized set one by one in the set order, and the duration of each display is the duration value associated with the order of the displayed model customization in the customized set; Based on the customized video, the matching video set is customized and adjusted.

[0054] After guiding the matching personnel to update the matching effect of the images to be matched based on the matching video set, the process also includes: When the target content of the update message associated with the matching requester's input in the matching video set is blank, the update message is decomposed into keywords to obtain multiple matching instructions; The matching instructions are combined to obtain multiple keyword combinations; among them, the type distribution of each matching instruction in the same keyword combination matches the standard type distribution, and the minimum number of matching parameters in the matching parameter association table of each matching instruction in the same keyword combination exceeds the threshold of the number of matching parameters. Each keyword combination is polled in turn. During each poll, the video is obtained based on the distribution of matching instructions in the polled keyword combinations. Matching parameter query conditions are obtained based on the matching instructions in the polled keyword combinations. The target matching parameters are then retrieved from the matching parameter library based on the matching parameter query conditions. After polling each keyword combination, the target matching parameters retrieved during each polling are combined to obtain the combined matching parameters; Based on the combined matching parameters, respond with an update message.

[0055] The step of combining the target matching parameters retrieved during each polling includes: Based on the multi-scale relationships between the pairwise target matching parameters retrieved during each polling, an updated combined model and an updated guiding video are created; The target matching parameters retrieved during each polling are associated with the updated combination model; when the matching personnel are guided by the update guidance video to update the matching effect based on the updated combination model after all target matching parameters are associated, the matching personnel can perceive all multi-scale relationships. The updated combined model, which associates all target matching parameters, and the updated guidance video are used as target matching parameters.

[0056] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.

[0057] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units described above can be implemented in hardware or as software functional units. The above are merely embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made based on the description and drawings of this application, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

[0058] The specific embodiments of the invention have been described in detail above, but they are only examples, and this application is not limited to the specific embodiments described above. For those skilled in the art, any equivalent modifications or substitutions to the invention are also within the scope of this application. Therefore, all equivalent changes, modifications, and improvements made without departing from the spirit and principles of this application should be covered within the scope of this application.

Claims

1. A multimodal fusion-based image and video matching system, characterized in that, include: The matching video unit is used to obtain a set of matching videos based on the matching effect data of the images to be matched associated with the matching tags when the matching personnel identify the matching tags on the images to be matched. The update unit is used to guide the matching personnel to update the matching effect of the images to be matched based on the set of matching videos. The update unit, based on the set of matching videos, guides users to update the matching results for the images to be matched, including: Based on the social network interest sets of people with matching needs, predict the matching effect of the image to be matched and recommend scenarios; Based on the matching effect of the image to be matched, a scene is recommended to obtain the recommended target scene; Based on the recommended target scenario, the matching video set is customized and adjusted; Guide users with matching needs to preview the customized and adjusted set of matching videos.

2. The image and video matching system based on multimodal fusion as described in claim 1, characterized in that, The update unit recommends scenes based on the matching effect of the image to be matched, and obtains the recommended target scene, including: The matching effect of the images to be matched is recommended in the scene by video frame representation, and multiple video frames are obtained; Frame graph analysis is performed on the correlation between the metadata of each video frame and between pairs of video frames to obtain frame graph matching features; Based on frame-matching features, key point responses are generated in the video, and multiple corresponding key points are generated in similar scenes. Based on the frame matching features, video frames are associated with each key point. Similar scenes obtained by associating each video frame with each key point are used as recommended target scenes.

3. The image and video matching system based on multimodal fusion as described in claim 1, characterized in that, The update unit customizes and adjusts the matching video set based on the recommended target scene, including: List the key points in the recommended target scene in descending order of the total number of associated video frames to obtain the key point set; Multiple partition subsets are determined from the set of key points; wherein the difference in the total number of video frames associated with adjacent key points in the same partition subset does not exceed the difference threshold or they belong to associated key points in the recommended target scene; Each partition subset is polled in turn. During each poll, the model is customized based on the video frames associated with each key point in the polled partition subset, and then associated with the polled partition subset. After polling each partition subset, the model customizations obtained in each polling are listed according to the order of the associated partition subsets in the keypoint set to obtain the customization set; The customized adjustment video is obtained; wherein, the customized adjustment video includes: matching the video set and displaying the customized models in the customized set one by one in the set order, and the duration of each display is the duration value associated with the order of the displayed model customization in the customized set; Based on the customized video, the matching video set is customized and adjusted.

4. The image and video matching system based on multimodal fusion as described in claim 1, characterized in that, The update unit, after guiding the matching personnel to update the matching effect of the images to be matched based on the matching video set, also includes: The deep update unit is used for: When the target content of the update message associated with the matching requester's input in the matching video set is blank, the update message is decomposed into keywords to obtain multiple matching instructions; The matching instructions are combined to obtain multiple keyword combinations; among them, the type distribution of each matching instruction in the same keyword combination matches the standard type distribution, and the minimum number of matching parameters in the matching parameter association table of each matching instruction in the same keyword combination exceeds the threshold of the number of matching parameters. Each keyword combination is polled in turn. During each poll, the video is obtained based on the distribution of matching instructions in the polled keyword combinations. Matching parameter query conditions are obtained based on the matching instructions in the polled keyword combinations. The target matching parameters are then retrieved from the matching parameter library based on the matching parameter query conditions. After polling each keyword combination, the target matching parameters retrieved during each polling are combined to obtain the combined matching parameters; Based on the combined matching parameters, respond with an update message.

5. The image and video matching system based on multimodal fusion as described in claim 4, characterized in that, The deep update unit performs matching parameter combinations on the target matching parameters retrieved during each polling, including: Based on the multi-scale relationships between the pairwise target matching parameters retrieved during each polling, an updated combined model and an updated guiding video are created; The target matching parameters retrieved during each polling are associated with the updated combination model; when the matching personnel are guided by the update guidance video to update the matching effect based on the updated combination model after all target matching parameters are associated, the matching personnel can perceive all multi-scale relationships. The updated combined model, which associates all target matching parameters, and the updated guidance video are used as target matching parameters.

6. A method for image and video matching based on multimodal fusion, characterized in that, include: When the matching personnel identify the matching tags on the images to be matched, they obtain a set of matching videos based on the matching effect data of the images to be matched associated with the matching tags. Based on the set of matching videos, guide the personnel who need to match to update the matching effect of the images to be matched; The step of guiding users to update the matching results for images based on the matching video set includes: Based on the social network interest sets of people with matching needs, predict the matching effect of the image to be matched and recommend scenarios; Based on the matching effect of the image to be matched, a scene is recommended to obtain the recommended target scene; Based on the recommended target scenario, the matching video set is customized and adjusted; Guide users with matching needs to preview the customized and adjusted set of matching videos.

7. The image and video matching method based on multimodal fusion as described in claim 6, characterized in that, The step of recommending scenarios based on the matching effect of the image to be matched, and obtaining the recommended target scenario, includes: The matching effect of the images to be matched is recommended in the scene by video frame representation, and multiple video frames are obtained; Frame graph analysis is performed on the correlation between the metadata of each video frame and between pairs of video frames to obtain frame graph matching features; Based on frame-matching features, key point responses are generated in the video, and multiple corresponding key points are generated in similar scenes. Based on the frame matching features, video frames are associated with each key point. Similar scenes obtained by associating each video frame with each key point are used as recommended target scenes.

8. The image and video matching method based on multimodal fusion as described in claim 6, characterized in that, The process of customizing and adjusting the matching video set based on the recommended target scene includes: List the key points in the recommended target scene in descending order of the total number of associated video frames to obtain the key point set; Multiple partition subsets are determined from the set of key points; wherein the difference in the total number of video frames associated with adjacent key points in the same partition subset does not exceed the difference threshold or they belong to associated key points in the recommended target scene; Each partition subset is polled in turn. During each poll, the model is customized based on the video frames associated with each key point in the polled partition subset, and then associated with the polled partition subset. After polling each partition subset, the model customizations obtained in each polling are listed according to the order of the associated partition subsets in the keypoint set to obtain the customization set; The customized adjustment video is obtained; wherein, the customized adjustment video includes: matching the video set and displaying the customized models in the customized set one by one in the set order, and the duration of each display is the duration value associated with the order of the displayed model customization in the customized set; Based on the customized video, the matching video set is customized and adjusted.

9. The image and video matching method based on multimodal fusion as described in claim 6, characterized in that, After guiding the matching personnel to update the matching effect of the images to be matched based on the matching video set, the process also includes: When the target content of the update message associated with the matching requester's input in the matching video set is blank, the update message is decomposed into keywords to obtain multiple matching instructions; The matching instructions are combined to obtain multiple keyword combinations; among them, the type distribution of each matching instruction in the same keyword combination matches the standard type distribution, and the minimum number of matching parameters in the matching parameter association table of each matching instruction in the same keyword combination exceeds the threshold of the number of matching parameters. Each keyword combination is polled in turn. During each poll, the video is obtained based on the distribution of matching instructions in the polled keyword combinations. Matching parameter query conditions are obtained based on the matching instructions in the polled keyword combinations. The target matching parameters are then retrieved from the matching parameter library based on the matching parameter query conditions. After polling each keyword combination, the target matching parameters retrieved during each polling are combined to obtain the combined matching parameters; Based on the combined matching parameters, respond with an update message.

10. The image and video matching method based on multimodal fusion as described in claim 9, characterized in that, The step of combining the target matching parameters retrieved during each polling includes: Based on the multi-scale relationships between the pairwise target matching parameters retrieved during each polling, an updated combined model and an updated guiding video are created; The target matching parameters retrieved during each polling are associated with the updated combination model; when the matching personnel are guided by the update guidance video to update the matching effect based on the updated combination model after all target matching parameters are associated, the matching personnel can perceive all multi-scale relationships. The updated combined model, which associates all target matching parameters, and the updated guidance video are used as target matching parameters.

Citation Information

Patent Citations

  • Visual rendering method and system for environmental art design

    CN120107446A