Label Determination Method, Device, Electronic Device, and Storage Medium
By obtaining the common characteristics of multimedia resources with social relationships with the target multimedia resources, and combining the label characteristics of preset labels, the target difference degree is calculated to determine the label, the problem of low accuracy of video tags in the prior art is solved, and more accurate label determination is achieved.
Patent Information
- Application Number
- CN202210266221.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-17
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-03-17
AI Technical Summary
When determining labels for videos in the prior art, the label accuracy is low based on the similarity between video features only.
By obtaining a multimedia resource with a social relationship with the target multimedia resource, determining its common characteristics relative to the target multimedia resource, and combining the label characteristics of the preset label, the target difference degree is calculated to determine the label.
Improve the accuracy of video tags, and ensure the accuracy and rationality of tags by considering the social attributes and differences in tag characteristics of multimedia resources.
Smart Images

Figure CN114639044B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of multimedia resource processing, and in particular, to a method, apparatus, electronic device, and storage medium for determining tags. Background Art
[0002] Tags can reflect the theme, content, etc. of multimedia resources, and play an important role in the collation and retrieval of multimedia resources. For example, each short video on a short video platform corresponds to its own tags, which can reflect the theme of the short video and the account interests, etc. Therefore, after an account publishes a short video, the short video platform needs to determine appropriate tags for the short video.
[0003] In the related art, the technical solution for determining tags for videos specifically includes: First, a certain number of videos are annotated with preset tags as the annotated videos. The video features of the annotated videos are extracted, and combined with the tags of the annotated videos to obtain the video features corresponding to each preset tag, so as to form a video tag index library. When adding tags to a new video, the video features of the new video are compared with the video features in the video tag index library to determine the video feature with the highest similarity to the video features of the new video, and the tag corresponding to the video feature is determined as the tag of the new video. However, the above method only tags videos based on the similarity between video features, with a single consideration factor, so the accuracy of the determined video tags may be relatively low. Summary of the Invention
[0004] The present disclosure provides a method, apparatus, electronic device, and storage medium for determining tags to at least solve the problem of relatively low accuracy of the determined tags in the related art. The technical solution of the present disclosure is as follows:
[0005] According to a first aspect of an embodiment of the present disclosure, a method for determining tags is provided, including: obtaining a first multimedia resource having a social relationship with a target multimedia resource, and determining a first common feature of the first multimedia resource relative to the target multimedia resource; the social relationship is used to represent that there is a social behavior between the accounts of different multimedia resources; determining the tag feature of a preset tag, and determining the target difference degree between the first common feature and the tag feature of the preset tag; in the case where the target difference degree is less than the preset difference degree, determining that the preset tag belongs to the target multimedia resource.
[0006] Optionally, determining the first common feature of the first multimedia resource relative to the target multimedia resource includes: inputting the resource features of the first multimedia resource and the resource features of the target multimedia resource into a first model pre-trained to output the first common feature; the first model is used to determine the common feature of multiple features according to the probability of similarity between multiple features.
[0007] Optionally, the above method further includes: obtaining a first sample multimedia resource having a social relationship with the target sample multimedia resource, and a sample label of the target sample multimedia resource; determining a label feature of the sample label, and using the resource feature of the target sample multimedia resource and the resource feature of the first sample multimedia resource as sample features, and using the label feature of the sample label as a supervision signal to train a preset first neural network to obtain a first model.
[0008] Optionally, obtaining a first sample multimedia resource having a social relationship with the target sample multimedia resource, and a sample label of the target sample multimedia resource includes: obtaining, from a pre-constructed heterogeneous graph, a first sample multimedia resource having a social relationship with the target sample multimedia resource, and a sample label of the target sample multimedia resource; the heterogeneous graph includes multiple sample multimedia resources, multiple sample labels, social relationships between any two sample multimedia resources, and attribution relationships between each sample multimedia resource and each sample label.
[0009] Optionally, determining a first common feature of the first multimedia resource relative to the target multimedia resource includes: determining the probability that the resource features of the first multimedia resources are similar to the resource feature of the target multimedia resource to obtain a similarity weight corresponding to each first multimedia resource; weighting the resource features of the first multimedia resources based on the similarity weights corresponding to the first multimedia resources to obtain a first common feature.
[0010] Optionally, the label feature of the preset label includes a label common feature, a video common feature, or a fusion feature obtained by fusing the label common feature and the video common feature; wherein, the label common feature includes the common feature of the same-class labels having the same category as the preset label relative to the preset label; the video common feature includes the common feature of the second multimedia resources having an attribution relationship with the preset label relative to the preset label.
[0011] Optionally, determining the label feature of the preset label includes: obtaining the same-class labels, and inputting the label features of the same-class labels and the label feature of the preset label into a pre-trained first model to output a label common feature; the first model is used to determine the common feature of multiple features according to the probability of similarity between the multiple features.
[0012] Optionally, determining the label feature of the preset label includes: obtaining a second multimedia resource, and inputting the resource feature of the second multimedia resource and the label feature of the preset label into a pre-trained first model to output a video common feature; the first model is used to determine the common feature of multiple features according to the probability of similarity between the multiple features.
[0013] According to a second aspect of the embodiments of the present disclosure, a tag determination device is provided, including an acquisition unit and a determination unit; the acquisition unit is configured to acquire a first multimedia resource having a social relationship with a target multimedia resource; the determination unit is configured to determine a first common feature of the first multimedia resource acquired by the acquisition unit relative to the target multimedia resource; the social relationship is used to represent that there are social behaviors between the accounts of different multimedia resources; the determination unit is further configured to determine the tag feature of a preset tag and determine the target difference degree between the first common feature and the tag feature of the preset tag; the determination unit is further configured to determine that the preset tag belongs to the target multimedia resource when the target difference degree is less than the preset difference degree.
[0014] Optionally, the determination unit is specifically configured to: input the resource feature of the first multimedia resource and the resource feature of the target multimedia resource into a pre-trained first model, and output the first common feature; the first model is configured to determine the common feature of multiple features according to the probability of similarity between multiple features.
[0015] Optionally, the acquisition unit is further configured to acquire a first sample multimedia resource having a social relationship with a target sample multimedia resource and the sample tag of the target sample multimedia resource; the determination unit is further configured to determine the tag feature of the sample tag; the tag determination device further includes a training unit, and the training unit is configured to use the resource feature of the target sample multimedia resource and the resource feature of the first sample multimedia resource as sample features, and use the tag feature of the sample tag as a supervision signal to train a preset first neural network to obtain the first model.
[0016] Optionally, the acquisition unit is specifically configured to: acquire a first sample multimedia resource having a social relationship with a target sample multimedia resource and the sample tag of the target sample multimedia resource from a pre-constructed heterogeneous graph; the heterogeneous graph includes multiple sample multimedia resources, multiple sample tags, the social relationship between any two sample multimedia resources, and the attribution relationship between each sample multimedia resource and each sample tag.
[0017] Optionally, the determination unit is specifically configured to: determine the probability of similarity of the resource features of each first multimedia resource relative to the resource feature of the target multimedia resource to obtain the similarity weight corresponding to each first multimedia resource; and based on the similarity weight corresponding to each first multimedia resource, weight the resource features of each first multimedia resource to obtain the first common feature.
[0018] Optionally, the tag features of the preset tag include tag common features, video common features, or fusion features obtained by fusing tag common features and video common features; among them, the tag common features include the common features of the same-category tags having the same category as the preset tag relative to the preset tag; the video common features include the common features of the second multimedia resource having an attribution relationship with the preset tag relative to the preset tag.
[0019] Optionally, the determining unit is specifically configured to: obtain the same-category tags, and input the tag features of the same-category tags and the tag features of the preset tag into a pre-trained first model, and output tag common features; the first model is used to determine the common features of multiple features according to the probability of similarity between multiple features.
[0020] Optionally, the determining unit is specifically configured to: obtain the second multimedia resource, and input the resource features of the second multimedia resource and the tag features of the preset tag into a pre-trained first model, and output video common features; the first model is used to determine the common features of multiple features according to the probability of similarity between multiple features.
[0021] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a processor, and a memory for storing instructions executable by the processor; wherein, the processor is configured to execute the instructions to implement the tag determination method as described in the first aspect above.
[0022] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which instructions are stored, and characterized in that when the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is enabled to execute the tag determination method as described in the first aspect above.
[0023] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, the computer program product includes computer instructions, and when the computer instructions are executed by the processor, the tag determination method as described in the first aspect above is implemented.
[0024] The technical solutions provided by the present disclosure at least bring the following beneficial effects: In the present disclosure, the tag determination device first obtains the first multimedia resource having a social relationship with the target multimedia resource, and determines the first common feature of the first multimedia resource relative to the target multimedia resource. Since the social relationship is used to represent that there are social behaviors between the accounts of different multimedia resources, there is a strong social attribute between the target multimedia resource and the first multimedia resource. Furthermore, this social attribute can be reflected by the first common feature of the first multimedia resource relative to the target multimedia resource, providing a basis for determining the tag of the target multimedia resource in the future. Further, the tag determination device determines the tag feature of the preset tag, and determines the target difference degree between the first common feature and the tag feature of the preset tag to measure whether the preset tag is suitable for the target multimedia resource. Furthermore, when the target difference degree is less than the preset difference degree, the tag determination device determines that the preset tag belongs to the target multimedia resource. Compared with the prior art that only considers the similarity between video features to label videos, the present disclosure combines the characteristics of strong social attributes of multimedia resources, extracts the common features of multimedia resources with social relationships, and uses the target difference degree between the common features and the tag features as the basis to determine the tags of multimedia resources. In this way, the determined tags are more accurate.
[0025] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Brief Description of the Drawings
[0026] The accompanying drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation to the present disclosure.
[0027] Figure 1 is a schematic diagram of a video tag shown according to an exemplary embodiment;
[0028] Figure 2 is a schematic diagram of the structure of a tag determination system shown according to an exemplary embodiment;
[0029] Figure 3 is a schematic diagram of the flowchart of a tag determination method shown according to an exemplary embodiment;
[0030] Figure 4 is a schematic diagram of the flowchart of a tag determination method shown according to an exemplary embodiment;
[0031] Figure 5 is a schematic diagram of the flowchart of a tag determination method shown according to an exemplary embodiment;
[0032] Figure 6It is the fourth flowchart diagram of a label determination method shown according to an exemplary embodiment;
[0033] Figure 7 It is a heterogeneous graph shown according to an exemplary embodiment;
[0034] Figure 8 It is the fifth flowchart diagram of a label determination method shown according to an exemplary embodiment;
[0035] Figure 9 It is the sixth flowchart diagram of a label determination method shown according to an exemplary embodiment;
[0036] Figure 10 It is a structural diagram of a label determination model shown according to an exemplary embodiment;
[0037] Figure 11 It is a structural diagram of a label determination device shown according to an exemplary embodiment;
[0038] Figure 12 It is a structural diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners
[0039] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0040] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described here can be implemented in an order other than those illustrated or described here. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0041] In addition, in the description of the embodiments of the present disclosure, unless otherwise specified, " / " means "or". For example, A / B can represent A or B. The "and / or" herein is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of the present disclosure, "a plurality" means two or more than two.
[0042] Before explaining the embodiments of the present disclosure in detail, some related technologies involved in the embodiments of the present disclosure will be introduced first.
[0043] In the embodiments of the present disclosure, the multimedia resources include but are not limited to videos, audios, pictures, texts, etc.
[0044] The tags can reflect the themes, contents, etc. of the multimedia resources, and play an important role in the collation and retrieval of the multimedia resources. For example, each short video corresponds to its own tags, and these tags can reflect the theme of the short video and the account interests, etc.
[0045] When determining tags for a short video, the existing method can extract tags from the short video title. Specifically, this method needs to first obtain the title of the short video, and segment the short video title to obtain the corresponding word sequence of the title. Further, part-of-speech tagging is performed on each word in the word sequence to generate the corresponding part-of-speech sequence of the word sequence. Finally, tags for the video are generated according to the word sequence and the part-of-speech sequence.
[0046] The existing method can also determine the tags of the short video based on the knowledge graph. Specifically, this method is based on the entity linking technology of the knowledge graph. According to the known knowledge graph, multiple candidate entities are extracted from the target video. Further, based on the pre-established video structured system, the knowledge graph, and the multiple candidate entities, the target main entity and / or target sub-entities corresponding to the target video are obtained; the vertical relationship between the main entity and the related sub-entities is defined in the video structured system. Finally, tags are assigned to the target video based on the main entity and / or the target sub-entities.
[0047] The existing method can also determine the tags of the short video by creating a video tag index library. Specifically, this method first uses preset tags to label a certain number of videos as the labeled videos. The video features of the labeled videos are extracted, and combined with the tags of the labeled videos to obtain the video features corresponding to each preset tag, so as to form a video tag index library. When adding tags to a new video, the video features of the new video are compared with the video features in the video tag index library, and the video feature with the highest similarity to the video features of the new video is determined, and the tag corresponding to this video feature is determined as the tag of the new video.
[0048] However, the above methods only label videos based on the similarity between video features, and the consideration factors are single. Therefore, the accuracy of the determined video tags may be relatively low. With the development of the Internet, the creation of existing multimedia resources usually has a strong social attribute. Such as Figure 1As shown, account B posted a short video about the new "weed-pulling cake". Follower account A of account B imitated account B and created a similar video. This phenomenon is called "behavior diffusion". The "behavior diffusion" phenomenon can lead to the same labels for videos corresponding to specific social relationships. Considering this feature, the present disclosure uses the social attributes of multimedia resources as an aid to determine labels for multimedia resources, thereby improving the accuracy of the determined video labels.
[0049] The label determination method provided by the embodiments of the present disclosure can be applied to a label determination system, which is used to solve the problem of low accuracy of the determined labels in the related art. Figure 2 A schematic structural diagram of the label determination system is shown. As Figure 2 shown, the label determination system 10 includes a label determination device 11 and an electronic device 12. The label determination device 11 is connected to the electronic device 12. The label determination device 11 and the electronic device 12 can be connected by a wired method or a wireless method. The embodiments of the present invention do not limit this.
[0050] The label determination device 11 is used to obtain a first multimedia resource having a social relationship with the target multimedia resource, and determine a first common feature of the first multimedia resource relative to the target multimedia resource. The label determination device 11 is also used to determine the label feature of a preset label, and determine the difference degree between the first common feature and the label feature of the preset label. The label determination device 11 is also used to determine whether the preset label belongs to the target multimedia resource according to the difference degree.
[0051] The label determination device 11 can be implemented in various electronic devices 12 that can process multimedia resources. Among them, the electronic device 12 can be a multimedia resource sharing platform, such as a short video sharing platform. The electronic device 12 at least has a multimedia resource storage device, a transmission device, and a multimedia resource playback device.
[0052] In different application scenarios, the label determination device 11 and the electronic device 12 can be independent devices or integrated into the same device. The embodiments of the present invention do not make specific limitations on this.
[0053] When the label determination device 11 and the electronic device 12 are integrated into the same device, the data transmission method between the label determination device 11 and the electronic device 12 is the data transmission between internal modules of the device. In this case, the data transmission process between the two is the same as "the data transmission process between the label determination device 11 and the electronic device 12 when they are independent of each other".
[0054] In the following embodiments provided by the embodiments of the present disclosure, the embodiments of the present disclosure take the label determination device 11 and the electronic device 12 being independently arranged as an example for description.
[0055] Figure 3 is a schematic flowchart of a tag determination method shown according to some exemplary embodiments. In some embodiments, the above tag determination method can be applied to a tag determination device, an electronic device as shown in Figure 1 and can also be applied to other similar devices.
[0056] As shown in Figure 3 , the tag determination method provided by the embodiments of the present disclosure includes the following S201-S206.
[0057] S201. The tag determination device obtains a first multimedia resource having a social relationship with the target multimedia resource.
[0058] Among them, the social relationship is used to represent that there are social behaviors between the accounts of different multimedia resources.
[0059] As a possible implementation manner, the tag determination device obtains a first multimedia resource having a social relationship with the target multimedia resource from the electronic device.
[0060] It should be noted that the social behavior includes interactive operations between accounts, such as a follow operation (account A follows account B), a like operation (account A likes the multimedia resource published by account B), and a notification operation (account A @ account B), etc. The embodiments of the present disclosure do not limit the specific social behavior.
[0061] Exemplarily, if account A of the target multimedia resource follows account B and account C, the tag determination device obtains the multimedia resources of account B and the multimedia resources of account C from the multimedia resource sharing platform.
[0062] S202. The tag determination device determines a first common feature of the first multimedia resource relative to the target multimedia resource.
[0063] As a possible implementation manner, the tag determination device performs conversion processing on the original data of the first multimedia resource to obtain the resource feature of the first multimedia resource, and performs conversion processing on the original data of the target multimedia resource to obtain the resource feature of the target multimedia resource; further, the tag determination device merges the resource feature of the first multimedia resource and the resource feature of the target multimedia resource to obtain the first common feature.
[0064] As another possible implementation manner, the tag determination device determines a first common feature of the first multimedia resource relative to the target multimedia resource according to a pre-trained feature determination model.
[0065] As another possible implementation, the tag determination device weights the resource features of each first multimedia resource according to the similarity weights of the resource features, and obtains a first common feature.
[0066] For the specific implementation of this step, reference can be made to the subsequent description of the embodiments of the present invention, and details will not be elaborated here.
[0067] S203. The tag determination device determines the tag features of the preset tag.
[0068] As a possible implementation, the tag determination device inputs the preset tag into a pre-trained feature determination model and outputs the tag features of the preset tag.
[0069] It should be noted that the preset tag is set in the tag determination device by the operation and maintenance personnel in advance. For example, the preset tag can be the tag of the historical video collected by the operation and maintenance personnel in advance.
[0070] Optionally, the tag features of the preset tag include tag common features, video common features, or fusion features obtained by fusing the tag common features and the video common features.
[0071] Among them, the tag common features include the common features of the same-category tags having the same category as the preset tag relative to the preset tag.
[0072] The video common features include the common features of the second multimedia resources having an attribution relationship with the preset tag relative to the preset tag.
[0073] For the specific implementation of this step, reference can be made to the subsequent description of the embodiments of the present invention, and details will not be elaborated here.
[0074] S204. The tag determination device determines the target difference degree between the first common feature and the tag features of the preset tag.
[0075] As a possible implementation, the tag determination device calculates the distance between the first common feature and the tag features of the preset tag according to a preset distance formula, and determines the calculated distance as the target difference degree.
[0076] It should be noted that the distance formula is set in the tag determination device by the operation and maintenance personnel in advance.
[0077] S205. The tag determination device determines whether the target difference degree is less than the preset difference degree.
[0078] As a possible implementation, the tag determination device compares the determined target difference degree with a preset standard difference degree to determine whether the target difference degree is less than the preset standard difference degree.
[0079] It should be noted that the standard difference degree is set in advance by the operation and maintenance personnel in the label determination device.
[0080] S206. When the target difference degree is less than the preset difference degree, the label determination device determines that the preset label belongs to the target multimedia resource.
[0081] As a possible implementation, if the target difference degree is less than or equal to the preset standard difference degree, the label determination device determines that the preset label belongs to the target multimedia resource; if the target difference degree is greater than the preset standard difference degree, the label determination device determines that the preset label does not belong to the target multimedia resource.
[0082] In some embodiments, the label determination device inputs the target difference degree into a preset scoring model and outputs a matching score. If the matching score is greater than the preset score, the label determination device determines that the preset label belongs to the target multimedia resource; if the matching score is less than or equal to the preset score, the label determination device determines that the preset label does not belong to the target multimedia resource.
[0083] It should be noted that the scoring model is set in advance by the operation and maintenance personnel in the label determination device. For example, the scoring model can be a Sigmoid model.
[0084] Exemplarily, the matching score s(v, t) of video v and label t = Sigmoid(h(t)(h(v)) T ), where h(t) represents the label feature, h(v) represents the video feature, and T represents the transpose of the video feature.
[0085] The technical solutions provided by the above embodiments at least bring the following beneficial effects: In the present disclosure, the tag determination device first obtains a first multimedia resource having a social relationship with the target multimedia resource, and determines a first common feature of the first multimedia resource relative to the target multimedia resource. Since the social relationship is used to represent that there are social behaviors between the accounts of different multimedia resources, there is a strong social attribute between the target multimedia resource and the first multimedia resource. Furthermore, this social attribute can be reflected by the first common feature of the first multimedia resource relative to the target multimedia resource, providing a basis for determining the tag of the target multimedia resource subsequently. Further, the tag determination device determines the tag feature of the preset tag, and determines the target difference degree between the first common feature and the tag feature of the preset tag to measure whether the preset tag is suitable for the target multimedia resource. Then, the tag determination device determines whether the preset tag belongs to the target multimedia resource according to the target difference degree. Compared with the prior art that only considers the similarity between video features to label videos, the present disclosure combines the characteristics of strong social attributes of multimedia resources, extracts the common features of multimedia resources with social relationships, and uses the target difference degree between the common feature and the tag feature as the basis to determine the tag of the multimedia resource. In this way, the determined tag is more accurate.
[0086] In one design, in order to determine the first common feature of the first multimedia resource relative to the target multimedia resource, as Figure 4 shown, the above S202 provided by the embodiments of the present disclosure specifically includes the following S2021-S2022:
[0087] S2021. The tag determination device determines the probability that the resource features of each first multimedia resource are similar to the resource features of the target multimedia resource, and obtains the similarity weight corresponding to each first multimedia resource.
[0088] As a possible implementation manner, the tag determination device first inputs the resource features of the target multimedia resource and the resource features of each first multimedia resource into a preset weight formula, determines the probability that the resource features of each first multimedia resource are similar to the resource features of the target multimedia resource, and obtains the similarity weight corresponding to each first multimedia resource.
[0089] It should be noted that the preset formula is preset by the operation and maintenance personnel in the tag determination device.
[0090] Exemplarily, the tag determination device first sets the type of the target multimedia resource v as a v , and the type of any first multimedia resource s ∈ N r (v) is a s , N r(v) represents the first multimedia resource having a social relationship with the target multimedia resource. The tag determination device calculates the query vector of the resource feature h(v) of the target multimedia resource v, the key vector of the resource feature h(s) of the first multimedia resource s, and the value vector of the resource feature h(s) of the first multimedia resource s according to the following preset formulas 1, 2, and 3: R d is the matrix corresponding to the social relationship r, with a dimension of d.
[0091]
[0092] Among them, Q, K, and V are preset initial vectors, h(v) is the resource feature of the target multimedia resource, and h(s) is the resource feature of the first multimedia resource. Both represent the linear transformation of the resource feature. represents the transformation matrix corresponding to the social relationship r.
[0093] Furthermore, the tag determination device calculates the probability that the resource features of each first multimedia resource are similar to the resource feature of the target multimedia resource according to preset formula 4:
[0094]
[0095] Among them, a r (v, s) represents the similarity weight corresponding to the first multimedia resource s. represents the transformation matrix corresponding to the social relationship r.
[0096] S2022. The tag determination device weights the resource features of each first multimedia resource based on the similarity weights corresponding to each first multimedia resource to obtain the first common feature.
[0097] As a possible implementation method, the tag determination device weights the resource features of each first multimedia resource according to a preset weighting formula based on the similarity weights corresponding to each first multimedia resource to obtain the first common feature.
[0098] Exemplarily, the tag determination device calculates the first common feature according to preset formula 5:
[0099]
[0100] Among them, m r (v) represents the first common feature. represents the value vector of any one first multimedia resource, and a r (v, s) represents the similarity weight corresponding to this first multimedia resource.
[0101] The technical solutions provided by the above embodiments at least bring the following beneficial effects: The label determination device determines the probability that the resource features of each first multimedia resource are similar to the resource features of the target multimedia resource, and obtains the similarity weights corresponding to each first multimedia resource, so as to evaluate the importance of each first multimedia resource. Further, the label determination device weights the resource features of each first multimedia resource based on the similarity weights corresponding to each first multimedia resource, and the obtained first common feature is more accurate.
[0102] In one design, in order to determine the first common feature of the first multimedia resource relative to the target multimedia resource, as Figure 5 shown, the above S202 provided by the embodiments of the present disclosure specifically includes the following S2023:
[0103] S2023. The label determination device inputs the resource features of the first multimedia resource and the resource features of the target multimedia resource into a pre-trained first model, and outputs the first common feature.
[0104] Among them, the first model is used to determine the common feature of multiple features according to the probability of similarity between multiple features.
[0105] In practical applications, after the label determination device inputs the resource features of the first multimedia resource and the resource features of the target multimedia resource into the pre-trained first model, the first layer in the first model will calculate the similarity weights corresponding to each first multimedia resource, and its calculation method can refer to the above S2021. The second layer in the first model will weight the resource features of each first multimedia resource based on the similarity weights corresponding to each first multimedia resource, and its calculation method can refer to the above S2022. Finally, the first model outputs the first common feature.
[0106] Exemplarily, the target multimedia resource is Video 1, and the first multimedia resources are Video 2 and Video 3. The label determination device inputs the video features of Video 1 (vector a 1 ), the video features of Video 2 (vector a 2 ), and the video features of Video 3 (vector a 3 ) into the pre-trained first model, and outputs the first common feature (vector a v ).
[0107] The technical solutions provided by the above embodiments at least bring the following beneficial effects: Since the first model can determine the common features of multiple features according to the similarity probability between the multiple features, after the label determination device inputs the resource features of the first multimedia resource and the resource features of the target multimedia resource into the pre-trained first model, the first model first calculates the similarity probability between the resource features of the first multimedia resource and the resource features of the target multimedia resource, and further outputs the first common feature according to the calculated probability. Therefore, it is more convenient and accurate to determine the first common feature through the first model.
[0108] In one design, in order to be able to train the first model, as Figure 6 shown, the label determination method provided by the embodiments of the present disclosure further includes the following S301-S303 before the above S2023:
[0109] S301. The label determination device obtains a first sample multimedia resource having a social relationship with the target sample multimedia resource, and the sample label of the target sample multimedia resource.
[0110] As a possible implementation manner, the label determination device obtains a first sample multimedia resource having a social relationship with the target sample multimedia resource, and the sample label of the target sample multimedia resource from the target sample set.
[0111] It should be noted that the target sample set includes the pre-collected first sample multimedia resource and the sample label of the target sample multimedia resource.
[0112] As another possible implementation manner, the label determination device obtains a first sample multimedia resource having a social relationship with the target sample multimedia resource, and the sample label of the target sample multimedia resource from the pre-constructed heterogeneous graph.
[0113] Among them, the heterogeneous graph includes multiple sample multimedia resources, multiple sample labels, the social relationship between any two sample multimedia resources, the attribution relationship between each sample multimedia resource and each sample label, and the category relationship between any two labels.
[0114] It should be noted that the heterogeneous graph is pre-constructed by the operation and maintenance personnel and stored in the label determination device.
[0115] Exemplarily, as Figure 7 shown, a representation form of a heterogeneous graph is shown, and the heterogeneous graph can be expressed as G(V, E). Among them, V is the set of nodes, which is composed of the sample multimedia resource set V 1 and the sample label set V 2 constitutes. E is the set of edges, which is composed of the is_subtopic_of (category relationship) relationship set E 1, the set E of has_tag (tag ownership relationship) relationships 2 and the set E of is_followed_by (social relationship) relationships 3 constitute. The is_followed_by relationship represents the social relationship between any two sample multimedia resources, pointing from the old multimedia resource to the new multimedia resource, that is, the old multimedia resource affects the new multimedia resource (for example, if account A follows account B, then the multimedia resources of account B are is_followed_by the multimedia resources of account A). The is_subtopic_of relationship represents the category relationship between any two tags (for example, strawberry cake and cheesecake are both in the cake category). The has_tag relationship represents the ownership relationship between each sample multimedia resource and each sample tag, coming from the already labeled video-tag data.
[0116] It can be understood that since the heterogeneous graph includes multiple sample multimedia resources, multiple sample tags, the social relationship between any two sample multimedia resources, the ownership relationship between each sample multimedia resource and each sample tag, and the category relationship between any two tags. Therefore, it is more efficient and convenient for the tag determination device to obtain the first sample multimedia resource that has a social relationship with the target sample multimedia resource and the sample tags of the target sample multimedia resource from the pre-constructed heterogeneous graph.
[0117] S302. The tag determination device determines the tag features of the sample tags.
[0118] As a possible implementation, the tag determination device converts the original data of the sample tags into vectors and uses the vectors of the sample tags as the tag features of the sample tags.
[0119] S303. The tag determination device uses the resource features of the target sample multimedia resource and the resource features of the first sample multimedia resource as sample features, and uses the tag features of the sample tags as the supervision signal to train the preset first neural network to obtain the first model.
[0120] As a possible implementation, the label determination device takes the resource features of the target sample multimedia resource and the resource features of the first sample multimedia resource as sample features, inputs them into a preset first neural network, and obtains a second common feature of the resource features of the first sample multimedia resource relative to the resource features of the target sample multimedia resource. The label determination device calculates the degree of difference between the second common feature and the label feature of the sample label. When the degree of difference between the second common feature and the label feature of the sample label is less than a preset threshold, the label determination device trains to obtain a first model. When the degree of difference between the second common feature and the label feature of the sample label is greater than or equal to the preset threshold, the label determination device uses a new target sample multimedia resource to iteratively train the first neural network until the degree of difference between the obtained second common feature and the label feature of the sample label is less than the preset threshold.
[0121] In practical applications, the first neural network can be any heterogeneous graph neural network. For example, a heterogeneous graph attention network (HGT), a hierarchical attention network (HAN).
[0122] Exemplarily, when the first neural network is an HGR, the label determination device takes the target sample multimedia resource as the central node and the first sample multimedia resource as the neighbor node, and inputs them into the HGR. Further, the label determination device takes the label feature of the sample label as a supervision signal to train the HGR until the degree of difference between the second common feature output by the HGR and the label feature of the sample label is less than the preset threshold.
[0123] The technical solution provided by the above embodiment at least brings the following beneficial effects: The label determination device first obtains the first sample multimedia resource having a social relationship with the target sample multimedia resource, and the sample label of the target sample multimedia resource; further, the label determination device determines the label feature of the sample label, and takes the resource features of the target sample multimedia resource and the resource features of the first sample multimedia resource as sample features, and takes the label feature of the sample label as a supervision signal to train a preset first neural network to obtain a first model. In this way, in the subsequent process, the label determination device can directly use the first model to determine the common feature of multiple features.
[0124] In one design, in order to determine the label feature of a preset label, as Figure 8 shown, S203 provided by the embodiments of the present disclosure specifically includes the following S2031-S2032:
[0125] S2031. The label determination device obtains homogeneous labels.
[0126] As a possible implementation, the label determination device obtains homogeneous labels that have the same category as the preset label from a pre-constructed heterogeneous graph.
[0127] S2032. The label determination device inputs the label features of the homogeneous labels and the label features of the preset label into a pre-trained first model, and outputs label common features.
[0128] Among them, the first model is used to determine the common features of multiple features according to the probability of similarity between the multiple features.
[0129] In practical applications, after the label determination device inputs the label features of the homogeneous labels and the label features of the preset label into the pre-trained first model, the first layer in the first model will calculate the similarity weights corresponding to each homogeneous label, and its calculation method can refer to the above S2021. The second layer in the first model will weight the resource features of each homogeneous label based on the similarity weights corresponding to each homogeneous label, and its calculation method can refer to the above S2022. Finally, the first model outputs label common features.
[0130] Exemplarily, the preset label is label 1, and the homogeneous labels are label 2 and label 3. The label determination device inputs the label features of label 1 (vector a1), the label features of label 2 (vector a2), and the label features of label 3 (vector a3) into the pre-trained first model, and outputs label common features (vector a t )
[0131] It should be noted that the training process of the first model here can refer to the above S301 - S303. The difference is that the label determination device uses the label features of the sample label and the label features of the sample homogeneous label as sample features, and uses the label features of the sample label as the supervision signal.
[0132] Exemplarily, when the first neural network is HGR, the label determination device inputs the sample label as the central node and the sample homogeneous label as the neighbor nodes into the HGR. Further, the label determination device uses the label features of the sample label as the supervision signal to train the HGR until the difference degree between the predicted label common features output by the HGR and the label features of the sample label is less than the preset threshold.
[0133] The technical solutions provided by the above embodiments at least bring the following beneficial effects: Since the first model can determine the common features of multiple features according to the probability of similarity between the multiple features, after the label determination device inputs the label features of the same-category labels and the label features of the preset label into the first model pre-trained in advance, the first model first calculates the probability of similarity between the label features of the preset label and the label features of the same-category labels, and further outputs the label common features according to the calculated probability. It can be seen that the label features determined by the first model can reflect the common features of the same-category labels that are specifically the same as the preset label, so that the output label features are more accurate.
[0134] In one design, in order to determine the label features of the preset label, as Figure 9 shown, the above S203 provided by the embodiments of the present disclosure specifically includes the following S2033-S2034:
[0135] S2033. The label determination device obtains the second multimedia resource.
[0136] As a possible implementation manner, the label determination device obtains the second multimedia resource having an attribution relationship with the preset label from the pre-constructed heterogeneous graph.
[0137] S2034. The label determination device inputs the resource features of the second multimedia resource and the label features of the preset label into the first model pre-trained in advance, and outputs the video common features.
[0138] Among them, the first model is used to determine the common features of multiple features according to the probability of similarity between the multiple features.
[0139] In practical applications, after the label determination device inputs the resource features of the second multimedia resource and the label features of the preset label into the first model pre-trained in advance, the first layer in the first model will calculate the similarity weights corresponding to the resource features of each second multimedia resource, and its calculation method can refer to the above S2021. The second layer in the first model will weight the resource features of each second multimedia resource based on the similarity weights corresponding to the resource features of each second multimedia resource, and its calculation method can refer to the above S2022. Finally, the first model outputs the video common features.
[0140] Exemplarily, the preset label is label 1, and the second multimedia resources are video 2 and video 3. The label determination device inputs the label features (vector a1) of label 1, the video features (vector a2) of video 2, and the video features (vector a3) of video 3 into the first model pre-trained in advance, and outputs the video common features (vector a v )
[0141] It should be noted that the training process of the first model here can refer to the above S301 - S303. The difference is that the label determination device uses the label feature of the sample label and the resource feature of the second sample multimedia resource as the sample feature, and uses the label feature of the sample label as the supervision signal.
[0142] Exemplarily, when the first neural network is HGR, the label determination device inputs the sample label as the central node and the second sample multimedia resource as the neighbor node into the HGR. Further, the label determination device uses the label feature of the sample label as the supervision signal to train the HGR until the difference degree between the predicted video common feature output by the HGR and the label feature of the sample label is less than the preset threshold.
[0143] The technical solutions provided by the above embodiments at least bring the following beneficial effects: Since the first model can determine the common feature of multiple features according to the similarity probability between multiple features, after the label determination device inputs the resource feature of the second multimedia resource and the label feature of the preset label into the pre - trained first model, the first model first calculates the similarity probability between the label feature of the preset label and the resource feature of the second multimedia resource, and further outputs the video common feature according to the calculated probability. It can be seen that the video feature determined by the first model can reflect the common feature of the second multimedia resource with respect to the preset label, which has a specific attribution relationship with the preset label. In this way, the output label feature is more accurate.
[0144] In one design, in order to determine the label feature of the preset label, the label determination device fuses the label common feature and the video common feature to obtain a fused feature.
[0145] Specifically, the label determination device fuses the label common feature and the video common feature according to the pre - trained second model to obtain a fused feature.
[0146] In practical applications, the second model can be any multi - modal aggregation model. For example, the multi - modal aggregation model can use the method of splicing plus linear transformation to fuse the label common feature and the video common feature:
[0147] h(t) = Linear r1 (m r1 (t)) + Linear r2 (m r2 (t))
[0148] Among them, h(t) represents the fused feature, and Linear r1 represents performing a linear transformation on the label common feature m r1 (t), and Linear r2 represents performing a linear transformation on the video common feature mr2 (t) performs a linear transformation.
[0149] In one design, to determine whether a preset label belongs to a target multimedia resource, a label determination device inputs the resource features of the target multimedia resource, the resource features of the first multimedia resource, and the label features of the preset label into a pre-trained label determination model, and outputs the target multimedia resource and a label that belongs to the target multimedia resource.
[0150] Exemplarily, as Figure 10 shown, a schematic structural diagram of a label determination model is shown. The label determination model is composed of three first models, one second model, and one judgment model. Among them, the three first models are respectively used to determine the first common feature h(v), the label common feature m r1 (t), and the video common feature m r2 (t). The second model is used to fuse the label common feature m r1 (t) and the video common feature m r2 (t) to obtain a fused feature h(t). The judgment model is used to calculate the similarity between h(v) and h(t), and output h(v) and h(t) whose similarity is greater than a preset threshold.
[0151] In one design, to train the label determination model, the label determination device uses the resource features of the target sample multimedia resource, the resource features of the first sample multimedia resource, and the label features of the sample label as sample features, and uses the sample label as a supervision signal to train the prediction label determination model. When the difference degree between the prediction label and the sample label is greater than a preset threshold, the parameters of the first model are adjusted, and the prediction label determination model is iteratively trained until the difference degree between the obtained prediction label and the sample label is less than or equal to the preset threshold.
[0152] The above embodiments mainly introduce the solutions provided by the embodiments of the present disclosure from the perspective of the device (equipment). It can be understood that, in order to implement the above methods, the device or equipment includes the corresponding hardware structure and / or software module for executing each method process, and these corresponding hardware structures and / or software modules for executing each method process can constitute a determination device for material information. Those skilled in the art should easily realize that, combined with the algorithm steps of each example described in the embodiments disclosed in this article, the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but this implementation should not be considered to exceed the scope of the present disclosure.
[0153] Embodiments of the present disclosure may divide functional modules for a device or equipment according to the above method examples. For example, the device or equipment may correspond to each function and divide each functional module, or integrate two or more functions into one processing module. The above integrated module may be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiments of the present disclosure is illustrative, merely a logical function division, and there may be other division methods in actual implementation.
[0154] Figure 11 is a schematic structural diagram of a tag determination device shown according to an exemplary embodiment. Referring to Figure 11 As shown, the tag determination device 40 provided by the embodiments of the present disclosure includes an acquisition unit 401 and a determination unit 402.
[0155] The acquisition unit 401 is configured to acquire a first multimedia resource having a social relationship with a target multimedia resource. For example, as Figure 3 shown, the acquisition unit 401 may be configured to execute S201.
[0156] The determination unit 402 is configured to determine a first common feature of the first multimedia resource acquired by the acquisition unit with respect to the target multimedia resource. The social relationship is used to characterize that there are social behaviors between the accounts of different multimedia resources. For example, as Figure 3 shown, the determination unit 402 may be configured to execute S202.
[0157] The determination unit 402 is further configured to determine the tag feature of a preset tag and determine the target difference degree between the first common feature and the tag feature of the preset tag. For example, as Figure 3 shown, the determination unit 402 may be configured to execute S203 - S204.
[0158] The determination unit 402 is further configured to, when the target difference degree is less than a preset difference degree, determine that the preset tag belongs to the target multimedia resource. For example, as Figure 3 shown, the determination unit 402 may be configured to execute S205 - S206.
[0159] Optionally, the determination unit 402 is specifically configured to: input the resource feature of the first multimedia resource and the resource feature of the target multimedia resource into a first model trained in advance, and output the first common feature. The first model is used to determine the common feature of multiple features according to the probability of similarity between multiple features.
[0160] Optionally, the acquisition unit 401 is further configured to acquire a first sample multimedia resource having a social relationship with the target sample multimedia resource, and the sample tag of the target sample multimedia resource.
[0161] The determination unit 402 is further configured to determine the label features of the sample label.
[0162] Optionally, the label determination device further includes a training unit 403. The training unit 403 is configured to use the resource features of the target sample multimedia resource and the resource features of the first sample multimedia resource as sample features, and use the label features of the sample label as a supervision signal to train a preset first neural network to obtain a first model.
[0163] Optionally, the obtaining unit 401 is specifically configured to: obtain, from a pre-constructed heterogeneous graph, a first sample multimedia resource having a social relationship with the target sample multimedia resource, and the sample label of the target sample multimedia resource. The heterogeneous graph includes a plurality of sample multimedia resources, a plurality of sample labels, the social relationship between any two sample multimedia resources, and the attribution relationship between each sample multimedia resource and each sample label.
[0164] Optionally, the determination unit 402 is specifically configured to: determine the probability that the resource features of each first multimedia resource are similar to the resource features of the target multimedia resource to obtain the similarity weight corresponding to each first multimedia resource. Based on the similarity weights corresponding to each first multimedia resource, weight the resource features of each first multimedia resource to obtain a first common feature.
[0165] Optionally, the label features of the preset label include label common features, video common features, or fusion features obtained by fusing the label common features and the video common features. Among them, the label common features include the common features of the same-class labels having the same category as the preset label relative to the preset label. The video common features include the common features of the second multimedia resources having an attribution relationship with the preset label relative to the preset label.
[0166] Optionally, the determination unit 402 is specifically configured to: obtain the same-class labels, and input the label features of the same-class labels and the label features of the preset label into a first model pre-trained to output label common features. The first model is used to determine the common features of multiple features according to the probability of similarity between the multiple features.
[0167] Optionally, the determination unit 402 is specifically configured to: obtain the second multimedia resources, and input the resource features of the second multimedia resources and the label features of the preset label into a first model pre-trained to output video common features. The first model is used to determine the common features of multiple features according to the probability of similarity between the multiple features.
[0168] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0169] Figure 12 This is a schematic structural diagram of an electronic device provided by the present disclosure. As Figure 12 , the electronic device 50 may include at least one processor 501 and a memory 502 for storing processor-executable instructions. Among them, the processor 501 is configured to execute the instructions in the memory 502 to implement the label determination method in the above embodiments.
[0170] In addition, the electronic device 50 may further include a communication bus 503 and at least one communication interface 504.
[0171] The processor 501 may be a central processing unit (CPU), a microprocessing unit, an ASIC, or one or more integrated circuits for controlling the execution of the program of the present disclosure solution.
[0172] The communication bus 503 may include a path for transmitting information between the above components.
[0173] The communication interface 504 uses any device such as a transceiver for communicating with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc.
[0174] The memory 502 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory may exist independently and be connected to the processing unit through a bus. The memory may also be integrated with the processing unit.
[0175] Among them, the memory 502 is used to store instructions for executing the solutions of the present disclosure and is controlled by the processor 501 for execution. The processor 501 is used to execute the instructions stored in the memory 502, thereby implementing the functions in the method of the present disclosure.
[0176] As an example, in combination with Figure 11 , the functions implemented by the acquisition unit 401, the determination unit 402, and the training unit 403 in the tag determination device 40 are the same as those of the processor 501 in Figure 12 .
[0177] In a specific implementation, as an embodiment, the processor 501 may include one or more CPUs, such as Figure 12 the CPU0 and CPU1 in
[0178] In a specific implementation, as an embodiment, the electronic device 50 may include multiple processors, such as Figure 12 the processor 501 and the processor 507 in
[0179] Each of these processors may be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. Here, the processor may refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).
[0180] Those skilled in the art can understand that Figure 12 the structure shown in
[0181] does not constitute a limitation on the electronic device 50 and may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component layout.
[0182] In addition, the present disclosure also provides a computer program product, including computer instructions, which, when running on an electronic device, cause the electronic device to execute the label determination method provided in the foregoing embodiments.
[0183] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the disclosure herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.
Claims
1. A method for determining a label, characterized in that, comprising: obtaining a first multimedia resource having a social relationship with a target multimedia resource, and determining a first common feature of the first multimedia resource relative to the target multimedia resource; the social relationship is used to characterize that there are social behaviors between the accounts of different multimedia resources; determining the label feature of the preset label according to the preset label and the first model, and determining the target difference degree between the first common feature and the label feature of the preset label; the label feature of the preset label includes a label common feature, a video common feature, or a fusion feature obtained by fusing the label common feature and the video common feature; wherein, the label common feature includes the common feature of the same-class labels having the same category as the preset label relative to the preset label; the video common feature includes the common feature of the second multimedia resource having an attribution relationship with the preset label relative to the preset label; the first model is used to determine the common feature of the multiple features according to the probability of similarity between the multiple features; in the case that the target difference degree is less than the preset difference degree, determining that the preset label belongs to the target multimedia resource.
2. The label determination method according to claim 1, characterized in that, the determining the first common feature of the first multimedia resource relative to the target multimedia resource includes: inputting the resource feature of the first multimedia resource and the resource feature of the target multimedia resource into the pre-trained first model, and outputting the first common feature.
3. The label determination method according to claim 2, characterized in that, the method further includes: obtaining a first sample multimedia resource having the social relationship with a target sample multimedia resource, and the sample label of the target sample multimedia resource; determining the label feature of the sample label, and using the resource feature of the target sample multimedia resource and the resource feature of the first sample multimedia resource as sample features, and using the label feature of the sample label as a supervision signal to train a preset first neural network to obtain the first model.
4. The label determination method according to claim 3, characterized in that, the obtaining a first sample multimedia resource having the social relationship with a target sample multimedia resource, and the sample label of the target sample multimedia resource includes: obtaining a first sample multimedia resource having the social relationship with a target sample multimedia resource, and the sample label of the target sample multimedia resource from a pre-constructed heterogeneous graph; the heterogeneous graph includes multiple sample multimedia resources, multiple sample labels, the social relationship between any two sample multimedia resources, and the attribution relationship between each sample multimedia resource and each sample label.
5. The label determination method according to claim 1, characterized in that, the determining the first common feature of the first multimedia resource relative to the target multimedia resource includes: Determine the probability that the resource characteristics of each of the first multimedia resources are similar to those of the target multimedia resource, and obtain the similarity weights corresponding to each of the first multimedia resources; Based on the similarity weights corresponding to each of the first multimedia resources, weight the resource characteristics of each of the first multimedia resources to obtain the first common characteristic.
6. The label determination method according to claim 1, characterized in that the determining the label characteristics of the preset label according to the preset label and the first model includes: obtaining the same-class labels, and inputting the label characteristics of the same-class labels and the label characteristics of the preset label into the pre-trained first model, and outputting the label common characteristic.
7. The label determination method according to claim 1, characterized in that the determining the label characteristics of the preset label according to the preset label and the first model includes: obtaining the second multimedia resource, and inputting the resource characteristics of the second multimedia resource and the label characteristics of the preset label into the pre-trained first model, and outputting the video common characteristic.
8. A label determination device, characterized in that it includes an acquisition unit and a determination unit; the acquisition unit is used to acquire the first multimedia resources having a social relationship with the target multimedia resource; the determination unit is used to determine the first common characteristic of the first multimedia resources acquired by the acquisition unit relative to the target multimedia resource; the social relationship is used to characterize that there are social behaviors between the accounts of different multimedia resources; the determination unit is further used to determine the label characteristics of the preset label according to the preset label and the first model, and determine the target difference degree between the first common characteristic and the label characteristics of the preset label; the label characteristics of the preset label include label common characteristics, video common characteristics or fusion characteristics obtained by fusing the label common characteristics and the video common characteristics; wherein, the label common characteristics include the common characteristics of the same-class labels having the same category as the preset label relative to the preset label; the video common characteristics include the common characteristics of the second multimedia resources having an attribution relationship with the preset label relative to the preset label; the first model is used to determine the common characteristics of multiple characteristics according to the probability of similarity between the multiple characteristics; the determination unit is further used to determine that the preset label belongs to the target multimedia resource when the target difference degree is less than the preset difference degree.
9. The label determination device according to claim 8, characterized in that the determination unit is specifically used for: inputting the resource characteristics of the first multimedia resource and the resource characteristics of the target multimedia resource into the pre-trained first model, and outputting the first common characteristic.
10. The label determination device according to claim 9, characterized in that the acquisition unit is further used to acquire the first sample multimedia resources having the social relationship with the target sample multimedia resource, and the sample label of the target sample multimedia resource; the determination unit is further used to determine the label characteristics of the sample label; The label determination device further includes a training unit, which is configured to use the resource features of the target sample multimedia resource and the resource features of the first sample multimedia resource as sample features, and use the label features of the sample label as a supervision signal to train a preset first neural network to obtain the first model.
11. The label determination device according to claim 10, wherein, the obtaining unit is specifically configured to: obtain, from a pre-constructed heterogeneous graph, a first sample multimedia resource having the social relationship with the target sample multimedia resource, and the sample label of the target sample multimedia resource; the heterogeneous graph includes a plurality of sample multimedia resources, a plurality of sample labels, the social relationship between any two sample multimedia resources, and the attribution relationship between each sample multimedia resource and each sample label.
12. The label determination device according to claim 8, wherein, the determining unit is specifically configured to: determine the probability that the resource features of each of the first multimedia resources are similar to the resource features of the target multimedia resource, to obtain a similarity weight corresponding to each of the first multimedia resources; weight the resource features of each of the first multimedia resources based on the similarity weights corresponding to each of the first multimedia resources, to obtain the first common feature.
13. The label determination device according to claim 8, wherein, the determining unit is specifically configured to: obtain the same category labels, and input the label features of the same category labels and the label features of the preset label into the pre-trained first model, and output the label common feature.
14. The label determination device according to claim 8, wherein, the determining unit is specifically configured to: obtain the second multimedia resource, and input the resource features of the second multimedia resource and the label features of the preset label into the pre-trained first model, and output the video common feature.
15. An electronic device, wherein, it includes: a processor and a memory for storing instructions executable by the processor; wherein, the processor is configured to execute the instructions to implement the label determination method according to any one of claims 1-7.
16. A computer-readable storage medium, on which instructions are stored, wherein, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the label determination method according to any one of claims 1-7.
17. A computer program product, wherein, the computer program product includes computer instructions, and when the computer instructions are executed by a processor, the label determination method according to any one of claims 1-7 is implemented.
Citation Information
Patent Citations
Label determination method, electronic equipment and storage medium
CN111400516A
Label prediction method and device, equipment and storage medium
CN113010705A