Video-based Object Evaluation Method, System, Device, and Storage Medium

Through the video-based object evaluation method, the matching degree between actors and character image characteristics is calculated using video-related text and image information, and the problem of low efficiency in actor recommendation in the prior art is solved, and fast and accurate actor recommendation is achieved.

CN114330963BActive Publication Date: 2025-06-27BEIJING IQIYI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111198504.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-14
Publication Date
2025-06-27
Estimated Expiration
2041-10-14

AI Technical Summary

Technical Problem

The prior art cannot quickly recommend suitable actors based on the emotional characteristics of the characters, resulting in staff spending a lot of time to screen actors who meet the characteristics of the characters.

Method used

Through the video-based object evaluation method, the image characteristics of the candidate objects and the target objects are determined using the video's associated text information and image information, the matching degree information is calculated, and the combination quality of the target objects in the video is evaluated.

Benefits of technology

It realizes the rapid and accurate identification of the emotional characteristics of characters in film and television works, improves the efficiency and accuracy of actor recommendations, and reduces the workload of manual screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114330963B_ABST
    Figure CN114330963B_ABST
Patent Text Reader

Abstract

The present application discloses a video-based object evaluation method, system, device, and storage medium. The method includes: determining at least two candidate objects included in the video and corresponding target objects based on the associated text information of the video; determining first matching degree information for each target object based on the image features of the candidate objects and the image features of the target objects; determining second matching degree information for each target object based on the image features of the target object and the image features of the corresponding reference object; and determining a combined quality evaluation result of at least two target objects in the video according to the second matching degree information and the first matching degree information of each target object, so that the interaction effect of the target objects as partners in the video can be determined through the combined quality evaluation result of the target objects in the video, creating a new evaluation method for evaluating video partner interaction, reducing the workload of manual selection, and saving human resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of video object evaluation, and particularly to a method, system, device, and storage medium for object evaluation based on video. Background Art

[0002] With the continuous progress of technology, information processing technology has been continuously developing. Information processing technology can be applied to many application fields. For example, information processing technology can be used to identify and locate target persons in videos; for another example, information processing technology can be used to build an intelligent casting system.

[0003] Among them, the intelligent casting system is a new commercial product provided by film and television developers. Based on information such as online video distribution, film industry, signed artists, self-media, and publicity and distribution teams in the film and television ecosystem, it constructs an integrated artist information release and screening platform, which improves the objectivity of decision-making in the casting process of film and television projects while improving the casting efficiency and quality of self-produced dramas by film and television developers. In addition, the intelligent casting system also provides a direct information exposure channel for including top stars, a large number of second- and third-tier artists, and acting newcomers, further narrowing the imbalance between the status and resources of artists in the circle, improving the efficiency of film production matchmaking, completing the effective utilization and allocation of resources, and realizing the maximization of value.

[0004] The film and television works in film and television projects tell stories through the language of the camera, and the emotions of characters are important narrative contents of the stories. When people watch film and television works, they can understand what the emotions of the characters are (that is, they need to understand what emotional state the characters in the film and television are in), why they are in this emotional state (that is, they need to understand the relationship between the emotional state and the circumstances of the characters), and how the emotions of such characters are (that is, they need to understand what impact this emotion will have on his and other characters' behaviors affected by him), so as to understand the stories happening in the film and television works and generate resonance, that is, to have spiritual communication with the video as the medium. This process is very complex, and the current exploration in various disciplines and technical fields is still relatively shallow. How to accurately and precisely identify the emotional characteristics of characters from the video content of film and television works to understand "how" the emotions of characters in film and television works are is a major problem in the existing technology.

[0005] Since there is currently a lack of an information processing method for determining the emotional characteristics of characters in film and television, it is impossible to quickly recommend corresponding actors for roles in film and television projects based on the emotional characteristics of characters. When staff members need to spend a lot of time selecting actors who meet the role characteristics from a large number of candidate actors based on the character information in film and television works, the workload of the staff members is increased. Summary of the Invention

[0006] The purpose of the embodiments of this application is to provide a video-based object evaluation method, system, device, and storage medium, so as to create a new evaluation method for evaluating video partner interaction and solve the problem in the prior art that it is impossible to quickly recommend corresponding actors for roles in film and television projects based on the emotional characteristics of characters.

[0007] In view of the above technical problems, this application is implemented through the following technical solutions:

[0008] In the first aspect of the implementation of this application, first, a video-based object evaluation method is provided, including: based on the associated text information of the video, determining at least two candidate objects included in the video, and the target objects corresponding to each of the candidate objects; according to the associated text information, determining the target relationship category of each candidate object and the image characteristics of each candidate object; according to the associated image information of each target object, determining the image characteristics of each target object; based on the image characteristics of the candidate objects and the image characteristics of the target objects, determining the first matching degree information of each target object; according to the target relationship category of each candidate object, selecting at least two reference objects corresponding to the target relationship category, and obtaining the image characteristics of each reference object; wherein, the reference objects and the target objects are in one-to-one correspondence; based on the image characteristics of the target objects and the image characteristics of their corresponding reference objects, determining the second matching degree information of each target object; according to the second matching degree information and the first matching degree information of each target object, determining the combined quality evaluation result of at least two of the target objects in the video.

[0009] Among them, according to the second matching degree information and the first matching degree information of each target object, determining the combined quality evaluation result of at least two of the target objects in the video includes: based on the first matching degree information and the second matching degree information of each target object, determining the target quality score corresponding to each target object; determining the object combination to which each target object belongs; performing weighted processing on the target quality scores corresponding to the target objects in the same object combination to obtain the combined quality evaluation result.

[0010] Among them, determining the second matching degree information of each target object based on the image features of the target object and the image features of the corresponding reference object includes: determining the edit distance between the image features of each target object and the reference image features, where the reference image features are the image features of the reference object corresponding to the target object; according to the edit distance between the image features of each target object and the reference image features, and the sample combination attributes corresponding to the reference image, determining the reference image feature fitting score corresponding to each target object; for each target object, determining the reference image feature fitting score as the second matching degree information.

[0011] Among them, determining the reference image feature fitting score corresponding to each target object according to the edit distance between the image features of each target object and the reference image features, and the sample combination attributes corresponding to the reference image includes: for each target object, determining the sample combination attributes corresponding to each reference object; if the sample combination attributes corresponding to the reference object are the first sample combination attributes, then determining the edit distance between the image features of the target object and the reference image features as the first reference image feature fitting score of the target object; if the sample combination attributes corresponding to the reference object are the second sample combination attributes, then determining the opposite number of the edit distance between the image features of the target object and the reference image features as the second reference image feature fitting score of the target object; for the same target object, adding the first reference image feature fitting score and the second reference image feature fitting score to obtain the reference image feature fitting score.

[0012] Among them, determining the first matching degree information of each target object based on the image features of the candidate object and the image features of the target object includes: determining the edit distance between the image features of each target object and the candidate image features, where the candidate image features are the image features of the candidate object corresponding to the target object; based on the edit distance between the image features of each target object and the candidate image features, determining the target image feature fitting score of each target object; for each target object, determining the target image feature fitting score as the first matching degree information.

[0013] Among them, determining the image features of each target object according to the associated image information of each target object includes: inputting the associated image information of each target object into a pre-trained feature interpretation model, where the associated image information is image information including at least two target objects; using the feature interpretation model to determine the image features corresponding to the target object in the associated image information.

[0014] Among them, according to the associated text information, determining the target relationship category of each candidate object includes: extracting keywords corresponding to the object relationship category from the associated text information; determining category score information according to the relationship category scores corresponding to the keywords; performing filtering processing according to the category score information to obtain the target relationship category.

[0015] Among them, the candidate object is a candidate role. Selecting at least two reference objects corresponding to the target relationship category according to the target relationship category of each candidate object includes: respectively searching for role attribute information corresponding to the target relationship category of each candidate object in a preset database; if the role attribute information matches the video role attribute information, then selecting the role object corresponding to the role attribute information as the reference object corresponding to the target relationship category.

[0016] Among them, the video role attribute information includes a target gender parameter and a target age parameter, and the role attribute information corresponding to the target relationship category includes a role gender parameter and a role age parameter; before selecting the role object corresponding to the role attribute information, it further includes: if the role gender parameter is the same as the target gender parameter, then determining an age deviation according to the role age parameter and the target age parameter; if the age deviation is within a preset deviation range, it is determined that the role attribute information corresponding to the target relationship category matches the video role attribute information.

[0017] In the second aspect of the implementation of the present application, a video-based object evaluation system is further provided, including:

[0018] An object determination module, configured to determine at least two candidate objects included in the video and target objects corresponding to each candidate object based on the associated text information of the video;

[0019] A category and feature determination module, configured to determine the target relationship category of each candidate object and the image features of each candidate object according to the associated text information;

[0020] A target image feature determination module, configured to determine the image features of each target object according to the associated image information of each target object;

[0021] A first matching degree determination module, configured to determine the first matching degree information of each target object based on the image features of the candidate object and the image features of the target object;

[0022] A reference object module, configured to select at least two reference objects corresponding to the target relationship category according to the target relationship categories of the candidate objects, and obtain the image features of each of the reference objects; wherein, the reference objects correspond to the target objects one by one;

[0023] A second matching degree determination module, configured to determine the second matching degree information of each target object based on the image features of the target object and the image features of the corresponding reference object;

[0024] An evaluation result determination module, configured to determine the combined quality evaluation result of at least two of the target objects in the video according to the second matching degree information and the first matching degree information of each target object.

[0025] In a third aspect of the implementation of the present application, an electronic device is further provided, including: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory complete communication with each other through the communication bus; the memory is used for storing a computer program; the processor is configured to implement the steps of the above-mentioned video-based object evaluation method when executing the program stored on the memory.

[0026] In a fourth aspect of the implementation of the present application, a computer-readable storage medium is further provided, and a computer program is stored in the computer-readable storage medium, and when the program is executed by a processor, the steps of the above-mentioned video-based object evaluation method are implemented.

[0027] The object evaluation method, system, device, and storage medium based on video provided by the embodiments of the present application determine at least two candidate objects included in the video and the target objects corresponding to each of the candidate objects through the associated text information based on the video; and determine the target relationship category of each of the candidate objects and the image features of each of the candidate objects according to the associated text information; and, determine the image features of each of the target objects according to the associated image information of each of the target objects; subsequently, based on the image features of the candidate objects and the image features of the target objects, determine the first matching degree information of each of the target objects; select at least two reference objects corresponding to the target relationship category according to the target relationship category of each of the candidate objects, and obtain the image features of each of the reference objects, so as to subsequently determine the second matching degree information of each target object based on the image features of the target object and the image features of its corresponding reference object, so that the combined quality evaluation result of at least two of the target objects in the video can be determined according to the second matching degree information and the first matching degree information of each target object, and the combined quality evaluation result can be used to evaluate the partner interaction effect of at least two target objects in the video, creating a new evaluation method for evaluating video partner interaction, being able to identify the image features of candidate objects from the video content of film and television works as accurately and precisely as possible, and being able to provide suitable object recommendations for film and television projects with different types of partner interactions, reducing the workload of manual selection and achieving the purpose of saving human resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art.

[0029] Figure 1 is the flowchart of the steps of an object evaluation method based on video provided by the embodiments of the present application;

[0030] Figure 2 is the flowchart of the steps of an object evaluation method based on video provided by an alternative embodiment of the present application;

[0031] Figure 3 is the structural block diagram of an object evaluation system based on video provided by the embodiments of the present application;

[0032] Figure 4 is the structural diagram of an electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0033] The following will describe the technical solutions of the embodiments of the present application with reference to the drawings in the embodiments of the present application.

[0034] According to an embodiment of the present application, a method for object evaluation based on video is provided. As Figure 1 shown, it is a flowchart of steps of a method for object evaluation based on video provided by an embodiment of the present application. The method for object evaluation based on video may specifically include the following steps:

[0035] Step S110: Based on the associated text information of the video, determine at least two candidate objects included in the video, and target objects corresponding to each of the candidate objects.

[0036] Among them, the associated text information of the video may refer to text information related to the video, such as movie project information. It should be noted that the movie project information may include various information of the movie project, such as the theme information of the movie project, the plot summary information, the character biography information, etc.; the theme information can be used to determine the theme of the movie project; the plot summary information can be used to determine the plot summary of the movie project; the character biography information can be used to determine the character biography in the movie project, such as some information describing age, character, and temperament of a person. The embodiments of the present application do not make specific limitations on this.

[0037] The candidate object may be an object to be selected in the video. The types of the candidate object may include but are not limited to characters and animals in a movie project. For example, the candidate object may be a fictional character in a movie. The target object corresponding to the candidate object may refer to a real object in real life used to play the candidate object in the video. For example, it may be an alternative actor in real life used to play a certain character in the video.

[0038] Step S120: Based on the associated text information, determine the target relationship category of each candidate object and the image characteristics of each candidate object.

[0039] The target relationship category of the candidate object may refer to the relationship category to which the candidate object belongs in the video. For example, it may be the partner category to which the candidate object belongs in the video, and this partner category can be determined according to the movie project information. Specifically, after obtaining the movie project information of a certain video, the obtained movie project image can be used as the associated text information of the video. Subsequently, a partner classification algorithm can be used to determine the partner category to which the candidate object belongs in the video according to the theme information, plot summary information, character biography information, etc. included in the movie project information, so as to use it as the target relationship category of the candidate object, so that the image characteristics of the corresponding reference object can be selected from the database according to this target relationship category later.

[0040] The image features of a candidate object can be a set of its own characteristic values shown by the candidate object in a video, and can be reflected in the keywords corresponding to each feature type of the candidate object in the video data. These keywords can be used to reflect the image features of the candidate object, such as temperament features, personality features, etc. In an alternative embodiment, the image features of the candidate object can be determined according to the character biography information used to describe the candidate object in the associated text information of the video. For example, in the case where the candidate object is a character role in a film and television project, the image features of the character role can be determined by the character biography information in the film and television project information. Further, the embodiments of the present application determine the image features of each candidate object based on the associated text information, which may specifically include: extracting the character biography information from the associated text information; determining the image features of each candidate object according to the character biography information. Among them, the character biography information may include text information that concisely describes the character features, and there will be keywords, phrases, etc. that express the character, such as information describing age, personality, and temperament of a person. The embodiments of the present application do not make specific limitations on this.

[0041] Step S130: Determine the image features of each target object based on the associated image information of each target object.

[0042] Specifically, for each target object, its corresponding image features can be determined through the keywords corresponding to the associated image information of the target object. It should be noted that the image features of the target object can be a set of its own characteristic values shown by the target object in the video roles and / or photo roles it has played in the past, and can be reflected in the keywords corresponding to each feature type in the video data and / or photo data. These keywords can be used to reflect the image features of the target object, such as being used to reflect the temperament features, personality features, etc. of the target object.

[0043] Step S140: Determine the first matching degree information of each target object based on the image features of the candidate object and the image features of the target object.

[0044] Specifically, after the embodiments of the present application determine the image features of the candidate object and the image features of the target object corresponding to the candidate object, for each target object, the image features of the target object and the image features of its corresponding candidate object can be used for calculation to obtain the feature fitting degree between the image features of each target object and its corresponding candidate object, and the feature fitting degree between the image features of each target object and its corresponding candidate object can be determined as the first matching degree information of each target object, so that the combined quality evaluation result of the target object in the video can be determined based on the first matching degree information subsequently.

[0045] Step S150: According to the target relationship categories of the respective candidate objects, select at least two reference objects corresponding to the target relationship category, and obtain the image features of each of the reference objects.

[0046] Among them, the reference objects correspond one by one to the target objects. Specifically, after the embodiments of the present application determine the target relationship category of the candidate object based on the associated text information of the video, according to the target relationship category, from a preset database, according to the role attribute information of the candidate object in the associated text information, select role objects corresponding to the role attribute information of each target object as being the same or similar, as the reference objects corresponding to the target object and / or the target relationship category, and the image features of each reference object can be extracted from the database. The database may store at least two role objects in one or more object combinations, the role attribute information and image features of each role object, and each object combination has a corresponding relationship category.

[0047] Step S160: Based on the image features of the target object and the image features of its corresponding reference object, determine the second matching degree information of each target object.

[0048] Specifically, after the embodiments of the present application obtain the image features of the reference object corresponding to the target relationship category, for each target object, the image features of the target object and the reference object can be used for calculation to obtain the feature fitting degree between the image features of each target object and the reference object. Subsequently, based on the feature fitting degree between the image features of the target object and the reference object, the second matching degree information of each target object can be determined.

[0049] Step S170: According to the second matching degree information and the first matching degree information of each target object, determine the combined quality evaluation result of at least two of the target objects in the video.

[0050] In a specific implementation, after the embodiments of the present application determine the second matching degree information and the first matching degree information of each target object, they can calculate based on the second matching degree information and the first matching degree information of each target object to obtain the target quality score corresponding to each target object. Subsequently, the target quality scores of the target objects in the same object combination can be weighted to obtain the quality score corresponding to the target objects in each object combination in the video, which is used as the combined quality evaluation result of the target objects in the video. Among them, each object combination can be composed of at least two target objects, and the target objects in each object combination are related to each other; the combined quality evaluation result of the target objects in the video can be used to represent the interactive effect of the film and television partners generated by the target objects as candidate objects in the video, creating a new evaluation method for evaluating the interactive effect of video partners, so that the interactive effect of the target objects in the video can be evaluated based on the combined quality evaluation result subsequently.

[0051] It can be seen that after the embodiments of the present application determine at least two candidate objects included in the video and the target objects corresponding to each candidate object based on the associated text information of the video, they can determine the first matching degree information of each target object by based on the image features of the candidate objects and the image features of the target objects, obtain the target relationship category and the image features of the candidate objects according to the obtained film and television project information, and can determine the second matching degree information of each target object based on the image features of the target object and the image features of its corresponding reference object. Subsequently, by relying on the second matching degree information and the first matching degree information of each target object, the combined quality evaluation result of at least two target objects in the video is determined, so that the interactive effect of the target objects in the video can be determined through the combined quality evaluation result of the target objects in the video, creating a new evaluation method for evaluating the interactive effect of video partners, converting the subjective human understanding and judgment into the interpretation of the computer, and being able to identify the emotional characteristics of candidate objects such as characters as accurately and precisely as possible from the video content of film and television works. Furthermore, suitable object recommendations can be provided for film and television projects with different types of partner interactions, reducing the workload of manual selection by staff and achieving the purpose of saving human resources.

[0052] In actual processing, the image features of an object can be determined by using a pre-trained feature interpretation model. Among them, the feature interpretation model can be a model obtained through pre-training, and specifically can be used to determine the image features shown by the object in a video or photo. Optionally, based on the above embodiments, the embodiments of the present application determine the image features of each target object according to the associated image information of each target object, which may specifically include: inputting the associated image information of each target object into a pre-trained feature interpretation model, where the associated image information is image information containing at least two target objects; using the feature interpretation model to determine the corresponding image features of the target object in the associated image information.

[0053] In a specific implementation, the associated image information of the target object may include multiple frames of associated images, and each frame of associated image may include at least two target objects in the same object combination; the feature interpretation model can process multiple frames of associated images containing at least two target objects to obtain the corresponding image features of each target object in the multiple frames of associated images as the image features of the target object, so as to subsequently determine the first matching degree information of the target object according to the image features of the target object; and the first matching degree information can be used to represent the degree of fit of the image features between the target object and the candidate object, and the degree of fit of the image features can be used to determine whether the target object is suitable for playing the role of the candidate object.

[0054] As an example of the present application, in the case of taking the roles in a film and television project as candidate objects and the alternative actors as target objects, the feature interpretation model can be used in advance to process the previous works and photos of the alternative actors as target objects to obtain the temperament feature vectors of the alternative actors, and the temperament feature vectors of the alternative actors can be used as the image features of the target object, so that the degree of fit of the image features between the alternative actors and the roles in the video can be determined according to the temperament feature vectors subsequently. Furthermore, based on the degree of fit of the image features, combined with the second matching degree information of the alternative actors, the combined quality evaluation result corresponding to the alternative actors in the video can be determined, so as to subsequently determine whether the alternative actors fit the roles in the video based on the combined quality evaluation result, achieving the purpose of accurately judging the quality of the partners and predicting the quality of the actor interactions, and avoiding the trouble for the staff to select actors who meet the role characteristics from a large number of candidate actors based on the character role information in the film and television works, thus reducing the work burden of the staff.

[0055] For example, associated images of two target objects that contain the same object combination can be obtained from video data. For instance, images in the video data can be sampled at a preset sampling interval, and object detection can be performed on the sampled images. Images that are detected to contain the two target objects in the object combination are used as associated images, or images that are detected to contain the two target objects and whose image quality meets the conditions are used as associated images. After obtaining the associated images of the target objects, an associated image sequence can be formed according to the playback position (i.e., playback time) of the associated images in the video data, as the associated image information of the target objects, so that subsequently, the multi-frame associated images containing the two target objects can be input into the feature interpretation model through this associated image series, and the feature interpretation model can process the multi-frame associated images containing the two target objects to obtain the corresponding image features of the target objects in the associated image information, which are used as the image features of the target objects.

[0056] It should be noted that the sampling interval can be an empirical value or an experimental value. For example, when the video data is 25 frames per second, the sampling interval can be set according to sampling 8 frames per second. Further, in this example, a standard image of the target object can be preset, and the label of this standard image can be the unique code of the target object. In the video data, starting from the first frame image, object detection can be performed on the first frame image, and after an interval of one sampling interval, object detection can be performed on the next frame image until all the video data has been detected. For the detected images, if the standard image of the target object is included in the image, then this image can be determined as the associated image of the target object, or if the standard image is included in the image and the image meets the image quality conditions, then this image is determined as the associated image of the target object. This example has no restrictions on this.

[0057] Specifically, after inputting multiple frames of associated images containing at least two target objects into a pre-trained feature interpretation model, the feature interpretation model can determine and output the corresponding image features of the target objects in the multiple frames of associated images as the image features of the target objects. Among them, the image features can be represented by an array or a vector. Each element position in the array can correspond to a feature type, and each dimension in the vector can correspond to a feature type. For example, when setting 15 dimensions in the image features, at least two keywords related to semantics can be set for each dimension, and keywords can be extracted from the character biography. It should be noted that the character biography is a text with concise words describing the character features, and there will be key words, phrases, etc. expressing the character. The key words, phrases, etc. can be extracted from the character biography through the bag-of-words model, and the invalid descriptions that do not describe the character and express negative meanings can be removed by using grammar structure knowledge, so as to obtain the keywords corresponding to the feature types. For example, when the dimensions in the image features include but are not limited to: dynamic-static, powerful-powerless, naive-mature, "dynamic", "static", "powerful", "powerless", "naive", "mature" are all keywords corresponding to the feature types.

[0058] Among them, the feature interpretation model can be trained through a pre-set training sample set. The training sample set can include multiple sample images. Each sample image includes an image of the target object and is pre-annotated with image features (i.e., annotated with multi-dimensional feature values). The image features are the real image features corresponding to the target object. When training the feature interpretation model, a supervised method can be used to train the feature interpretation model. The input of the feature interpretation model can be a single sample image, and the output can be the multi-dimensional feature values corresponding to the single sample image. Each dimension in the multi-dimension corresponds to a feature type. The multi-dimensional feature values are the predicted image features. Using the real multi-dimensional feature values of the input sample image, it is determined whether the feature interpretation model converges. After the feature interpretation model converges, the training of the feature interpretation model is stopped.

[0059] Furthermore, during the process of training the feature interpretation model, the sub-models corresponding to each dimension can be trained independently. Each dimension's sub-model requires multiple sample images (such as about 100,000) for independent training so that the sub-model can accurately determine the image features of the target object in each sample image. Subsequently, aggregation processing can be performed on the image features corresponding to the target object in each frame of the associated images to determine the image features corresponding to the target object in the multiple frames of associated images as the image features of the target object.

[0060] Among them, each sub-model of the feature interpretation model can output an operation value of one dimension. The bottom layer of the feature interpretation model discretizes the operation value of each dimension separately, and discretizes the operation values of each dimension into multiple polarity values. For example, in the case where the initial feature interpretation model is an image regression model including 15 sub-models, the 15 sub-models are independently trained to obtain a feature interpretation model capable of determining personality characteristics.

[0061] It should be noted that the polarity values in this example can include -1, 0, and 1. The range of the operation values output by the sub-model of each dimension is [-1, 1]. The bottom layer of the feature interpretation model discretizes each operation value in a preset manner, so that each operation value becomes one of the three polarity values of -1, 0, and 1. Further, a threshold can be used to discretize the operation value. For example, the first threshold is 0.5, the second threshold is -0.5. When the operation value is greater than -0.5 and less than 0.5, the polarity value is 0. When the operation value is greater than or equal to 0.5, the polarity value is 1. When the operation value is less than or equal to -0.5, the polarity value is -1. In the case where each feature type corresponds to two antonyms, -1 indicates that the target object belongs to the first feature word. 1 indicates that the target object belongs to the second feature word, and the first feature word and the second feature word are antonyms of each other. 0 indicates that the target object belongs neither to the first feature word nor to the second feature word, and is between the first feature word and the second feature word. In this way, if 15 dimensions are set in the image feature, then the dimension values of the 15 dimensions are connected together to form a 15-dimensional image feature, which can theoretically represent 3 to the 15th power of personality characteristics. In the case where each feature type corresponds to one feature word, -1 indicates a first-level match, 0 indicates a second-level match, and 1 indicates a third-level match.

[0062] After the feature interpretation model is trained, the feature interpretation model can be used to determine the image feature corresponding to the target object in multiple frames of associated images, that is, the feature interpretation model can be used to determine the image feature of the target object, so that subsequently, by comparing the image feature of the target object with the image feature of the candidate object in the film and television project, the video partner interaction effect of the target object in the film and television project can be determined.

[0063] Of course, the feature interpretation model can also be used to determine the image feature corresponding to the candidate object in the video, and the embodiments of the present application do not limit this. Among them, the feature types involved in the image feature can be determined according to user needs, and the embodiments of the present application do not make specific limitations on this either.

[0064] In addition, embodiments of the present application can also determine the image characteristics of candidate objects according to the character biography information in the associated text information of the video, so that the image feature matching degree between the candidate object and the candidate object in the video can be determined based on the image characteristics subsequently, and the partner quality result of the candidate object corresponding to the video can be determined according to the image feature matching degree, so that it can be determined whether the candidate object plays the candidate object in the video based on the partner quality result. For example, when the candidate object is a character in a film and television project, the image and temperament characteristics of the character can be used as the image characteristics of the candidate object. After extracting the character biography information from the film and television project information, the image and temperament characteristics of the character can be determined according to the character biography information, so that subsequently, based on the image and temperament characteristics of the character, using a feature interpretation model, the image and temperament matching degree score between the candidate object and the character of the film and television project can be determined. Furthermore, the partner quality of the candidate object in the film and television project can be determined based on the image and temperament matching degree score between the candidate object and the character of the film and television project, and the interaction effect of the candidate object can be predicted to determine whether the candidate object is suitable to play the role of the film and television project.

[0065] Referring to Figure 2 , a flowchart showing the steps of a method for object evaluation based on video provided by an optional embodiment of the present application is shown. As Figure 2 shown, the method for object evaluation based on video may include the following steps:

[0066] Step S210, based on the associated text information of the video, determine at least two candidate objects included in the video, and the target objects corresponding to each of the candidate objects.

[0067] Step S220, according to the associated text information, determine the target relationship category of each candidate object and the image characteristics of each candidate object.

[0068] As an example of the present application, the film and television project information of a video can be used as the associated text information of the video, so as to determine the target relationship category of the roles as candidate objects in the video and the image characteristics of each role. Specifically, after obtaining the film and television project information of the video, the film and television project information can be used, such as the keywords related to the materials showing the plot development and character relationships in the film and television project information, such as the theme, plot summary, character biography, etc., to judge which object relationship category the roles of each film and television partner in the film and television project belong to, and the relationship category scores corresponding to the keywords can be given according to the pre-established partner interaction keyword value system, so that the object relationship categories in the film and television project and the category scores corresponding to each object relationship category can be determined based on the relationship category scores corresponding to the keywords, and then the object relationship categories with too low category scores can be filtered out to determine the object relationship category with a higher category score as the target relationship category. For example, the object relationship category with the highest category score can be used as the target relationship category of the film and television project.

[0069] Furthermore, the embodiment of the present application determines the target relationship category of each of the candidate objects according to the associated text information, which may specifically include the following sub-steps:

[0070] Step S2201: Extract the keywords corresponding to the object relationship category from the associated text information;

[0071] Step S2202: Determine the category score information according to the relationship category scores corresponding to the keywords;

[0072] Step S2203: Perform filtering processing according to the category score information to obtain the target relationship category.

[0073] In actual processing, a numerical system of partner interaction keywords can be established in advance. For example, keywords can be determined based on information such as the displayed partner interaction methods, the feelings between characters, and the personalities of characters, and corresponding relationship category scores can be given to the keywords. Each keyword can correspond to multiple object relationship categories, and each partner type can be assigned corresponding scores. After extracting keywords from the associated text information of the video, the scores of the same category included in the keywords can be added up to obtain the category scores corresponding to each object relationship category as category score information. Subsequently, based on this category score information, each object relationship category can be filtered to delete the object relationship categories with relatively low category scores, so that the object relationship categories with relatively high category scores can be determined as the target relationship categories corresponding to the film and television project information. For example, when the category score exceeds a certain score threshold, the object relationship category corresponding to the category score can be determined as the target relationship category. If the category scores of multiple partners all exceed a certain score threshold, it can be considered that the partner roles of this film and television project belong to these several object relationship categories at the same time, that is, the target relationship category includes these several object relationship categories, thus solving the problem that the partner role of a certain film and television project has the characteristics of two or more categories at the same time. Another example is that through keyword score calculation, the object relationship category with the highest category score can be determined as the target relationship category corresponding to this film and television partner role. The embodiments of the present application do not make specific limitations on this.

[0074] Step S230: Input the associated image information of each of the target objects into a pre-trained feature interpretation model, where the associated image information is image information including at least two target objects.

[0075] Step S240: Use the feature interpretation model to determine the image features corresponding to each target object in the associated image information.

[0076] Step S250: Based on the image features of the candidate objects and the image features of the target objects, determine the first matching degree information of each target object.

[0077] Furthermore, in the embodiments of the present application, based on the image features of the candidate objects and the image features of the target objects, determining the first matching degree information of each target object may specifically include: determining the edit distance between the image features of each target object and the candidate image features, where the candidate image features are the image features of the candidate objects corresponding to the target objects; based on the edit distance between the image features of each target object and the candidate image features, determining the target image feature fitting score of each target object; for each target object, determining the target image feature fitting score as the first matching degree information.

[0078] For example, when the target object is a candidate actor, the number of candidate actors is multiple, and the candidate object is the video protagonist in a film and television project, after determining the image features of each candidate actor, the edit distance between the image features of the candidate actor and the image features of the video object can be used to determine the image feature fitting score corresponding to the candidate actor, so as to be used as the target image feature fitting score of the candidate actor, so that it can be determined whether the image of the candidate actor is suitable for playing the video role based on the target image feature fitting score of the candidate actor in the subsequent process. Among them, the target image feature fitting score of the candidate actor can be used as the first matching degree information, which specifically refers to the fitting score between the image features of the video role and the image features of the candidate actor, and is used to represent the image fitting degree between the video role and the candidate actor.

[0079] Specifically, when the temperament feature vector of the candidate actor is used as the image feature of the candidate actor, after determining the image feature vector of the video role based on the character biography information in the film and television project information, the edit distance between the temperament feature vector of each candidate actor and the image feature vector of the video role can be calculated through the temperament feature vector of each candidate actor and the image feature vector of the video role, and the image feature fitting score corresponding to each candidate actor can be determined respectively based on the edit distance between the temperament feature vector of each candidate actor and the image feature vector of the video role. For example, the edit distance between the temperament feature vector of the candidate actor and the image feature vector of the video role can be used as the target image feature fitting score of the candidate actor, and then the target image feature fitting score can be used as the first matching degree information of the candidate actor, so that it can be determined the total partner quality score corresponding to the candidate actor based on the target image feature fitting score of the candidate actor in the subsequent process.

[0080] Step S260, select at least two reference objects corresponding to the target relationship category according to the target relationship category of each candidate object, and obtain the image features of each reference object.

[0081] Among them, the reference object corresponds to the target object one by one. For example, when the candidate object is a fictional character in the film and television project information and the target relationship category is the partner category to which the to-be-selected character belongs in the video, after determining the partner category to which the character belongs in the video, the partner example object that matches the attribute information and has the same partner type can be selected from the preset database according to the attribute information of the fictional character in the film and television project information as the reference object, and the temperament characteristics of the partner example object can be obtained as the image characteristics of the reference image. It should be noted that the temperament characteristics of the partner example object can include the temperament characteristics of the successful example object and / or the temperament characteristics of the failed example object; the temperament characteristics of the successful example object can represent the characteristic temperament of the object in the successful partner combination example determined by the classic partner model; the characteristic temperament of the failed example object can represent the characteristic temperament of the object in the failed partner combination example determined by the classic partner model.

[0082] In actual processing, different temperament partner combinations can be summarized from the previous film and television works by artificial professional knowledge to establish an expert knowledge base. For example, 15 different temperament partner combinations can be summarized, etc.; and according to the quality of the chemical effect generated by the role interaction, successful example objects and failed example objects can be listed for each partner combination. Subsequently, based on the expert knowledge base, through a pre-trained feature interpretation model, such as a pre-trained temperament feature interpretation model, the temperament vectors of the two types of example objects of different partner combinations can be obtained, and the temperament characteristics can be extracted from the temperament vectors of the two types of example objects of different partner combinations through the classic partner model to obtain the characteristic temperament of the successful example object and the characteristic temperament of the failed example object, and the characteristic temperament of the successful example object and the characteristic temperament of the failed example object can be stored in the database as the temperament characteristics of the partner example object corresponding to the partner category for subsequent partner evaluation.

[0083] Furthermore, when the candidate object in the embodiment of the present application is the candidate role in the video, the above-mentioned selection of at least two reference objects corresponding to the target relationship category according to each candidate object specifically includes: respectively searching for the role attribute information corresponding to the target relationship category of each candidate object in the preset database; if the role attribute information matches the video role attribute information, then select the role object corresponding to the role attribute information as the reference object corresponding to the target relationship category. Among them, the video role attribute information refers to the attribute information of the candidate role.

[0084] Specifically, after determining the target relationship category of each candidate role, based on the target relationship category, role objects corresponding to the object relationship category identical to the target relationship category can be found in the database, and the role attribute information corresponding to the role objects can be compared with the video role attribute information to determine whether the role attribute information corresponding to the role objects is the same as or similar to the video role attribute information. Thus, when the role attribute information corresponding to the role objects is the same as or similar to the video role attribute information, it can be determined that the role attribute information corresponding to the role objects matches the video role attribute information. Furthermore, the role objects can be selected as the reference objects corresponding to the target relationship category, so that subsequently, based on the image characteristics of the reference objects corresponding to the target relationship category, it can be determined whether the target object is suitable for playing the candidate role in the video.

[0085] In an alternative embodiment, the video role attribute information in the embodiments of the present application may include a target gender parameter and a target age parameter. Among them, the target gender parameter can represent the gender of the candidate role, and the target age parameter can represent the age of the candidate role. Further, the role attribute information corresponding to the target relationship category in the embodiments of the present application may include a role gender parameter and a role age parameter. Before selecting the role object corresponding to the role attribute information, it may further include: if the role gender parameter is the same as the target gender parameter, then based on the role age parameter and the target age parameter, the age deviation is determined; if the age deviation is within the preset deviation range, it is determined that the role attribute information corresponding to the target relationship category matches the video role attribute information.

[0086] Specifically, when the role gender parameter of a certain role object corresponding to the target relationship category is the same as the target gender parameter, the age parameter corresponding to the role object can be subtracted from the target age parameter to obtain the age deviation corresponding to the role object; subsequently, by determining whether the age deviation corresponding to the role object is within the preset deviation range, it can be determined whether the role attribute information corresponding to the role object matches the video role attribute information. If the age deviation corresponding to the candidate feature is within the preset deviation range, it can be determined that the role attribute information corresponding to the role object matches the video role attribute information, and further, the role object can be determined as the reference object corresponding to the target relationship category; if the age deviation corresponding to the role object is not within the preset deviation range, it can be determined that the role attribute information corresponding to the role object does not match the video role attribute information, and further, the role object can be ignored and not selected as the reference object corresponding to the target relationship category.

[0087] As an example of the present application, after determining the partner category to which the candidate character belongs based on the obtained film and television project information, the characteristic temperament vectors of the classic successful partner combination example objects and the characteristic temperament vectors of the classic failed partner combination example objects of the same partner type with the same gender and similar age can be selected from the database according to the age and gender of the candidate characters in the film and television project information, so as to serve as the image characteristics of the reference objects corresponding to the candidate characters. For example, in the case where the image characteristics of the reference object corresponding to the candidate character include the temperament vector of the successful example object and / or the temperament vector of the failed example object, the characteristic temperament vector of the selected classic successful partner combination example can be determined as the temperament vector of the successful example object, and the characteristic temperament vector of the classic failed partner combination example can be determined as the temperament vector of the failed example object.

[0088] Step S270, based on the image characteristics of the target object and the image characteristics of the corresponding reference object, determine the second matching degree information of each target object.

[0089] Optionally, the embodiment of the present application determines the second matching degree information of each target object based on the image characteristics of the target object and the image characteristics of the corresponding reference object, which may specifically include: determining the edit distance between the image characteristics of each target object and the reference image characteristics, where the reference image characteristics are the image characteristics of the reference object corresponding to the target object; according to the edit distance between the image characteristics of each target object and the reference image characteristics, and the sample combination attributes corresponding to the reference image, determine the reference image characteristic fitting score corresponding to each target object; for each target object, determine the reference image characteristic fitting score as the second matching degree information.

[0090] As an example of the present application, in combination with the above example, after obtaining the image characteristics of the reference object corresponding to the video character, the image characteristics of the reference object can be used as the reference image characteristics, so as to determine the reference image characteristic fitting score corresponding to each candidate actor according to the image characteristics of the reference object and the personality characteristics of the candidate actor, as the second matching degree information of the target object. Among them, the reference image characteristic fitting score as the second matching degree information may refer to the fitting score between the image characteristics of the reference object and the image characteristics of the candidate actor, and can specifically be used to represent the image characteristic fitting degree between the reference object corresponding to the video character and the candidate actor, so that it can be determined whether the candidate actor is suitable to play the video character according to the image characteristic fitting degree between the reference object and the candidate actor.

[0091] For example, when the image feature of the reference object corresponding to the video character is the temperament feature vector (0, -1, 0, 0, 0, 1, 1, 1, 0, 1, 0, 0), if the image feature of candidate actor A is the temperament vector (0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 0, 0), then the edit distance between the temperament feature vector of the video character and the temperament vector (0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 0, 0) of candidate actor A is calculated. For example, each eigenvalue in the temperament feature vector of the video character is operated on with each eigenvalue in the temperament vector (0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 0, 0) of candidate actor A, and then the result values corresponding to each eigenvalue after the operation are accumulated to obtain the edit distance between the temperament feature vector of the video character and the temperament vector (0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 0, 0) of candidate actor A. For example, the edit distance between the temperament feature vector of the video character and the temperament vector of candidate actor A is 0 + 1 + 0 + 0 + 0 + 0 + 0 + 0 + 0 + 1 + 0 + 0 = 2, which is used as the second matching degree information of candidate actor A; if the image feature of candidate actor B is the temperament vector (0, 1, 0, -1, 0, 1, 0, 0, -1, -1, -1, 0), and the edit distance between the temperament feature vector of the video character and the temperament vector (0, 1, 0, -1, 0, 1, 0, 0, -1, -1, -1, 0) of candidate actor B is 0 + 1000 + 0 + 1 + 0 + 0 + 1 + 1 + 1 + 1000 + 1 + 0 = 2005, then the edit distance corresponding to candidate actor A is less than the edit distance corresponding to candidate actor B. Therefore, compared with candidate actor B, candidate actor A is more suitable for the image of a cold young talent like the video character. If the pre-set edit distance threshold for "image fitting" is 5, then the edit distance 2 corresponding to candidate actor A is less than the edit distance threshold 5, so candidate actor A can be characterized as "fitting this type of role"; while the edit distance 2005 corresponding to candidate actor B is greater than the edit distance threshold 5, and candidate actor B can be characterized as "not fitting this type of role".

[0092] Furthermore, the embodiment of the present application determines the fitting score of the reference image feature corresponding to each target object according to the edit distance between the image feature of each target object and the reference image feature, and the sample combination attribute corresponding to the reference image, which may include: for each target object, determining the sample combination attribute corresponding to each reference object; if the sample combination attribute corresponding to the reference object is the first sample combination attribute, then determining the edit distance between the image feature of the target object and the reference image feature as the first reference image feature fitting score of the target object; if the sample combination attribute corresponding to the reference object is the second sample combination attribute, then determining the negative value of the edit distance between the image feature of the target object and the reference image feature as the second reference image feature fitting score of the target object; for the same target object, adding the first reference image feature fitting score and the second reference image feature fitting score to obtain the reference image feature fitting score. Among them, the sample combination attribute corresponding to the reference object may refer to the partner example attribute of the image feature of the reference object, such as determining the sample combination attribute corresponding to the reference object according to whether the image feature corresponding to the reference object is a classic successful example or a classic failed example.

[0093] In actual processing, the sample combination attribute corresponding to the reference object can be divided into a first sample combination attribute and a second sample combination attribute; the first sample combination attribute can represent the positive sample combination attribute corresponding to the successful example; the second sample combination attribute can represent the negative sample combination attribute corresponding to the failed example. Specifically, when the sample combination attribute corresponding to the reference object is the first sample combination attribute, for example, when it is determined according to the sample combination attribute corresponding to the image feature of the reference object that the image feature corresponding to the reference object is a classic successful example, the original value of the edit distance between the image feature of the reference object and the personality feature of the candidate actor can be retained as the first reference image feature fitting score of the candidate actor; when the sample combination attribute corresponding to the reference object is the second sample combination attribute, for example, when it is determined according to the relationship quality attribute corresponding to the image feature of the reference object that the image feature corresponding to the reference object is a classic failed example, the negative value of the edit distance between the image feature of the reference object and the personality feature of the candidate actor can be taken as the second reference image feature fitting score of the candidate actor. Subsequently, for the same candidate actor, the first reference image feature fitting score and the second reference image feature fitting score of the candidate actor can be added to obtain the reference image feature fitting score of the candidate actor as the second matching degree information of the candidate actor.

[0094] Step S280, determining the combined quality evaluation result of at least two of the target objects in the video according to the second matching degree information and the first matching degree information of each target object.

[0095] Further, the embodiment of the present application determines the combined quality evaluation result of at least two of the target objects in the video according to the second matching degree information and the first matching degree information of each target object, which may specifically include: determining the target quality score corresponding to each target object based on the first matching degree information and the second matching degree information of each target object; determining the object combination to which each target object belongs; and performing weighted processing on the target quality scores corresponding to the target objects in the same object combination to obtain the combined quality evaluation result of each target object in the video in the object combination.

[0096] For example, when determining the fitting score of the target image feature and the reference image feature of the candidate actor, the fitting score of the target image feature and the fitting score of the reference image feature can be added together to determine the total fitting score between the candidate actor and the video character, so that the total fitting score can be used as the target quality score corresponding to the candidate actor. Subsequently, the partner combination of the candidate actor can be used as the object combination to which the candidate actor belongs, and by statistically calculating the target quality scores corresponding to all candidate actors in the same partner combination, the combined quality evaluation result corresponding to the candidate actors in the video in the partner combination can be determined, so that appropriate candidate actors can be selected based on the combined quality evaluation result for the candidate objects subsequently, realizing intelligent casting, avoiding the trouble for the staff to spend a lot of time screening for actors who meet the character characteristics from a large number of candidate actors based on the character information in the film and television works, and reducing the work burden of the staff.

[0097] It can be seen that after the embodiments of the present application determine at least two candidate objects included in the video and the target objects corresponding to each candidate object based on the associated text information of the video, the target relationship category of each candidate object and the image features of each candidate object can be determined according to the associated text information, and the image features of each target object can be determined according to the associated image information of each target object. Therefore, based on the image features of the candidate objects and the image features of the target objects, the first matching degree information of each target object can be determined, and at least two reference objects corresponding to the target relationship category can be selected according to the target relationship category of each candidate object, so as to determine the second matching degree information of each target object according to the image features of the reference objects, and then the candidate objects can be selected. Furthermore, the target quality score corresponding to each target object can be determined according to the second matching degree information and the first matching degree information of each target object. Subsequently, the target quality scores corresponding to the target objects in the same object combination can be weighted to obtain the combined quality evaluation result of each target object in the video in each object combination, so that suitable target objects can be selected as candidate objects in the video based on the combined quality evaluation result. Among them, the first matching degree information can represent the feature fitting degree between the target object and the candidate object, and can specifically be used to measure the personality similarity between the target object and the candidate object to determine whether the target object is used as a candidate object in a film and television project.

[0098] In actual processing, the video-based object evaluation method of the embodiments of the present application can have many application scenarios. For example, it can be applied in the scenario of intelligent casting to provide suitable actor recommendations for film and television projects with different types of partner interactions; or it can be applied in project evaluation and prediction systems to estimate the market prospects of the casting plan for a project. The embodiments of the present application do not make specific limitations on this.

[0099] Specifically, in the intelligent casting process of this embodiment, starting from the basic principle of casting, a feature interpretation model can be used to interpret the works of candidate actors using an algorithm, and general features of the candidate actors (i.e., the image features of the candidate actors) can be obtained therefrom. That is, the subjective character image temperament is quantified. It is also possible to determine the target relationship category (i.e., the partner category in the video) and the image features of the candidate objects (i.e., the image features of the characters in the video) based on the film and television project information. That is, based on the film and television project information, a computer algorithm is used to understand the human interaction in the video, determine the partner category in the video and the image features of the characters, and can use the algorithm to identify the emotional features of the characters as accurately and precisely as possible from the video content of the film and television works, convert the subjective human judgment into the interpretation of the computer, effectively judge the actor temperament and the quality of partner interaction, and realize providing an algorithm interpretation for the chemical reaction generated by the interaction between people, thereby being able to provide a new idea based on partner interaction for the casting of film and television projects, saving the casting time and improving the work efficiency of casting.

[0100] By applying the embodiments of the present application to the intelligent casting scenario, an integrated artist information release and screening platform can be constructed in aspects such as online video distribution, film industry, signed artists, self-media, publicity and distribution teams, etc., improving the objectivity of decision-making in the casting process and the casting efficiency and quality. At the same time, it can also provide a direct information batch channel for top stars, a large number of second- and third-tier artists, and acting newcomers, further narrowing the imbalance between artist status and resources, improving the efficiency of film shooting matchmaking, completing the effective utilization and allocation of resources, and realizing the maximization of value.

[0101] The embodiments of the present application also provide a video-based object evaluation system. As Figure 3 shown, the video-based object evaluation system provided by the embodiments of the present application specifically includes the following modules:

[0102] An object determination module 310, configured to determine at least two candidate objects included in the video and corresponding target objects based on the associated text information of the video;

[0103] A category and feature determination module 320, configured to determine the target relationship category of each candidate object and the image features of each candidate object based on the associated text information;

[0104] A target image feature determination module 330, configured to determine the image features of each target object based on the associated image information of each target object;

[0105] A first matching degree determination module 340, configured to determine the first matching degree information of each target object based on the image features of the candidate objects and the image features of the target objects;

[0106] A reference object module 350 is configured to select at least two reference objects corresponding to the target relationship category according to the target relationship category of each candidate object, and obtain the image features of each reference object; wherein, the reference object corresponds to the target object one by one;

[0107] A second matching degree determination module 360 is configured to determine the second matching degree information of each target object based on the image features of the target object and the image features of the corresponding reference object;

[0108] An evaluation result determination module 370 is configured to determine the combined quality evaluation result of at least two target objects in the video according to the second matching degree information and the first matching degree information of each target object.

[0109] Wherein, the evaluation result determination module 370 may include the following sub-modules:

[0110] A target quality scoring sub-module is configured to determine the target quality score corresponding to each target object based on the first matching degree information and the second matching degree information of each target object;

[0111] An object combination determination sub-module is configured to determine the object combination to which each target object belongs;

[0112] An evaluation result sub-module is configured to perform weighted processing on the target quality scores corresponding to the target objects in the same object combination to obtain the combined quality evaluation result.

[0113] Wherein, the second matching degree determination module 360 may include the following sub-modules:

[0114] An edit distance sub-module is configured to determine the edit distance between the image features of each target object and the reference image features, and the reference image features are the image features of the reference object corresponding to the target object;

[0115] A reference image feature fitting score sub-module is configured to determine the reference image feature fitting score corresponding to each target object according to the edit distance between the image features of each target object and the reference image features, and the sample combination attributes corresponding to the reference image;

[0116] A second matching degree information sub-module is configured to determine the reference image feature fitting score as the second matching degree information for each target object.

[0117] Among them, the reference image feature fitting sub-module is specifically configured to: for each target object, determine the sample combination attributes corresponding to each reference object; if the sample combination attributes corresponding to the reference object are the first sample combination attributes, then determine the edit distance between the image features of the target object and the reference image features as the first reference image feature fitting score of the target object; if the sample combination attributes corresponding to the reference object are the second sample combination attributes, then determine the negative value of the edit distance between the image features of the target object and the reference image features as the second reference image feature fitting score of the target object; for the same target object, add the first reference image feature fitting score and the second reference image feature fitting score to obtain the reference image feature fitting score.

[0118] Among them, the first matching degree determination module 340 includes the following sub-modules:

[0119] The edit distance determination sub-module is used to determine the edit distance between the image features of each target object and the candidate image features, and the candidate image features are the image features of the candidate object corresponding to the target object;

[0120] The target image feature fitting sub-module is used to determine the target image feature fitting score of each target object based on the edit distance between the image features of each target object and the candidate image features;

[0121] The first matching degree information sub-module is used to, for each target object, determine the target image feature fitting score as the first matching degree information.

[0122] Among them, the target image feature determination module 330 includes the following sub-modules:

[0123] The input sub-module is used to input the associated image information of each target object into a feature interpretation model pre-trained, and the associated image information is image information including at least two target objects;

[0124] The image feature determination sub-module is used to use the feature interpretation model to determine the image features corresponding to the target object in the associated image information.

[0125] Among them, the category and feature determination module 320 includes the following sub-modules:

[0126] The keyword extraction sub-module is used to extract keywords corresponding to the object relationship category from the associated text information;

[0127] The category score determination sub-module is used to determine category score information according to the relationship category scores corresponding to the keywords;

[0128] A target relationship category sub-module, configured to perform filtering processing according to the category score information to obtain the target relationship category.

[0129] Wherein, the candidate object is a candidate role, and the reference object module 350 includes the following sub-modules:

[0130] A search sub-module, configured to search for role attribute information corresponding to the target relationship category of each candidate object in a preset database respectively;

[0131] A role object selection sub-module, configured to select a role object corresponding to the role attribute information as a reference object corresponding to the target relationship category when the role attribute information matches the video role attribute information.

[0132] Wherein, the video role attribute information includes a target gender parameter and a target age parameter, and the role attribute information corresponding to the target relationship category includes a role gender parameter and a role age parameter; the reference object module 350 further includes: a matching sub-module. The matching sub-module is configured to determine an age deviation according to the role age parameter and the target age parameter when the role gender parameter is the same as the target gender parameter; if the age deviation is within a preset deviation range, it is determined that the role attribute information corresponding to the target relationship category matches the video role attribute information.

[0133] The functions of the system described in the embodiments of the present application have been described in the above method embodiments. Therefore, for the details not described in the description of the embodiments of the present application, reference may be made to the relevant descriptions in the foregoing embodiments, and details will not be repeated here.

[0134] The embodiments of the present application further provide an electronic device, such as Figure 4 shown, including a processor 410, a communication interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communication interface 420, and the memory 430 complete communication with each other through the communication bus 440.

[0135] The memory 430 is used to store a computer program.

[0136] The processor 410, when executing the program stored on the memory 430, implements the steps of the method for object evaluation based on video as described in any of the above embodiments.

[0137] Exemplarily, the processor 410, when executing the program stored on the memory 430, implements the following steps:

[0138] Based on the video-associated text information, determine at least two candidate objects included in the video, and the target objects corresponding to each of the candidate objects;

[0139] According to the associated text information, determine the target relationship category of each candidate object and the image features of each candidate object;

[0140] According to the associated image information of each target object, determine the image features of each target object;

[0141] Based on the image features of the candidate objects and the image features of the target objects, determine the first matching degree information of each target object;

[0142] According to the target relationship category of each candidate object, select at least two reference objects corresponding to the target relationship category, and obtain the image features of each reference object; wherein, the reference objects correspond to the target objects one by one;

[0143] Based on the image features of the target objects and the image features of their corresponding reference objects, determine the second matching degree information of each target object;

[0144] According to the second matching degree information and the first matching degree information of each target object, determine the combined quality assessment result of at least two of the target objects in the video.

[0145] The communication bus mentioned in the above terminal may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0146] The communication interface is used for communication between the above terminal and other devices.

[0147] The memory may include a Random Access Memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0148] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0149] In another embodiment provided by the present application, a computer-readable storage medium is further provided. The computer-readable storage medium stores a computer program, and when the program is executed by a processor, it implements the steps of the object evaluation method based on video described in any one of the above embodiments.

[0150] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that a computer can access, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a Solid State Disk (SSD)).

[0151] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.

[0152] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the relevant content.

[0153] The above description is only a preferred embodiment of the present application and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application are included in the protection scope of the present application.

Claims

1. A video-based object evaluation method, characterized in that, Including: Based on the associated text information of the video, determining at least two candidate objects included in the video, and target objects corresponding to each of the candidate objects, where the candidate objects are fictional character roles in the video, and the target objects are real objects used to play the candidate objects in the video; Determining the target relationship category of each of the candidate objects and the image features of each of the candidate objects according to the associated text information; Determining the image features of each of the target objects according to the associated image information of each of the target objects, including: inputting the associated image information of each of the target objects into a pre-trained feature interpretation model to determine the image features of the target objects; Based on the image features of the candidate objects and the image features of the target objects, determining the first matching degree information of each of the target objects; Selecting at least two reference objects corresponding to the target relationship category according to the target relationship category of each of the candidate objects, and obtaining the image features of each of the reference objects; where the reference objects correspond one-to-one with the target objects, and the reference objects are role objects corresponding to the target objects with the same or similar role attribute information; Based on the image features of the target objects and the image features of their corresponding reference objects, determining the second matching degree information of each target object; Determining the combined quality evaluation result of at least two of the target objects in the video according to the second matching degree information and the first matching degree information of each target object, including: determining the target quality score corresponding to each target object based on the first matching degree information and the second matching degree information of each target object; determining the combined quality evaluation result based on the target quality score.

2. The video-based object evaluation method according to claim 1, wherein Determining the combined quality evaluation result of at least two of the target objects in the video according to the second matching degree information and the first matching degree information of each target object, including: Determining the target quality score corresponding to each target object based on the first matching degree information and the second matching degree information of each target object; Determining the object combination to which each target object belongs; Performing a weighted process on the target quality scores corresponding to the target objects in the same object combination to obtain the combined quality evaluation result.

3. The video-based object evaluation method according to claim 1 or 2, characterized in that, The determining the second matching degree information of each target object based on the image features of the target objects and the image features of their corresponding reference objects includes: Determining the edit distance between the image features of each target object and the reference image features, where the reference image features are the image features of the reference object corresponding to the target object; Determining the reference image feature fitting score corresponding to each target object according to the edit distance between the image features of each target object and the reference image features, and the sample combination attribute corresponding to the reference image; For each target object, determining the reference image feature fitting score as the second matching degree information.

4. The video-based object evaluation method according to claim 3, wherein Determining the fitting score of the reference image feature corresponding to each target object according to the edit distance between the image feature of each target object and the reference image feature, and the sample combination attribute corresponding to the reference image, includes: For each target object, determining the sample combination attribute corresponding to each reference object; If the sample combination attribute corresponding to the reference object is the first sample combination attribute, determining the edit distance between the image feature of the target object and the reference image feature as the first fitting score of the reference image feature of the target object; If the sample combination attribute corresponding to the reference object is the second sample combination attribute, determining the opposite number of the edit distance between the image feature of the target object and the reference image feature as the second fitting score of the reference image feature of the target object; For the same target object, adding the first fitting score of the reference image feature and the second fitting score of the reference image feature to obtain the fitting score of the reference image feature.

5. The video-based object evaluation method according to claim 1 or 2, characterized in that, Determining the first matching degree information of each target object based on the image feature of the candidate object and the image feature of the target object, includes: Determining the edit distance between the image feature of each target object and the candidate image feature, where the candidate image feature is the image feature of the candidate object corresponding to the target object; Based on the edit distance between the image feature of each target object and the candidate image feature, determining the fitting score of the target image feature of each target object; For each target object, determining the fitting score of the target image feature as the first matching degree information.

6. The video-based object evaluation method according to claim 1 or 2, characterized in that Determining the image feature of each target object according to the associated image information of each target object, includes: Inputting the associated image information of each target object into a pre-trained feature interpretation model, where the associated image information is image information including at least two target objects; Using the feature interpretation model to determine the image feature corresponding to the target object in the associated image information.

7. The video-based object evaluation method according to claim 1 or 2, characterized in that, Determining the target relationship category of each candidate object according to the associated text information, includes: Extracting keywords corresponding to the object relationship category from the associated text information; Determining the category score information according to the relationship category score corresponding to the keyword; Performing filtering processing according to the category score information to obtain the target relationship category.

8. The video-based object evaluation method according to claim 1, characterized in that The candidate object is a candidate role, and selecting at least two reference objects corresponding to the target relationship category according to the target relationship category of each candidate object, includes: In a preset database, respectively searching for the role attribute information corresponding to the target relationship category of each candidate object; If the role attribute information matches the video role attribute information, selecting the role object corresponding to the role attribute information as the reference object corresponding to the target relationship category.

9. The video-based object evaluation method according to claim 8, wherein The video role attribute information includes a target gender parameter and a target age parameter, and the role attribute information corresponding to the target relationship category includes a role gender parameter and a role age parameter; Before selecting the character object corresponding to the character attribute information, it further includes: If the character gender parameter is the same as the target gender parameter, determine the age deviation based on the character age parameter and the target age parameter; If the age deviation is within the preset deviation range, determine that the character attribute information corresponding to the target relationship category matches the video character attribute information.

10. A video-based object evaluation system, characterized in that, It includes: An object determination module, configured to determine at least two candidate objects included in the video and target objects corresponding to each candidate object based on the associated text information of the video, where the candidate object is a fictional character in the video, and the target object is a real object used to play the candidate object in the video in real life; A category and feature determination module, configured to determine the target relationship category of each candidate object and the image features of each candidate object according to the associated text information; A target image feature determination module, configured to determine the image features of each target object according to the associated image information of each target object, including: inputting the associated image information of each target object into a pre-trained feature interpretation model to determine the image features of the target object; A first matching degree determination module, configured to determine the first matching degree information of each target object based on the image features of the candidate object and the image features of the target object; A reference object module, configured to select at least two reference objects corresponding to the target relationship category according to the target relationship category of each candidate object, and obtain the image features of each reference object; where the reference object corresponds to the target object one by one, and the reference object is a character object corresponding to the same or similar character attribute information as the target object; A second matching degree determination module, configured to determine the second matching degree information of each target object based on the image features of the target object and the image features of its corresponding reference object; An evaluation result determination module, configured to determine the combined quality evaluation result of at least two target objects in the video according to the second matching degree information and the first matching degree information of each target object, including: determining the target quality score corresponding to each target object based on the first matching degree information and the second matching degree information of each target object; determining the combined quality evaluation result based on the target quality score.

11. An electronic device, characterized in that, It includes: A processor, a communication interface, a memory, and a communication bus, where the processor, the communication interface, and the memory complete communication with each other through the communication bus; The memory is used to store a computer program; The processor is configured to implement the steps of the object evaluation method based on video according to any one of claims 1-9 when executing the program stored on the memory.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the object evaluation method based on video according to any one of claims 1-9.

Citation Information

Patent Citations

  • Generic mapping for tracking target object in video sequence

    US20170132472A1

  • Text information matching degree detection method and apparatus, computer device and storage medium

    WO2020258506A1