Character description generation method and device for media resource data, equipment and medium
By mapping different modal features of media resource data to the same vector space and generating text descriptions using a text-trained decoder, the complexity problem caused by modal differences is solved, and efficient text description generation of media resource data is achieved.
Patent Information
- Application Number
- CN202311438931.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-31
- Publication Date
- 2025-05-02
AI Technical Summary
There is a problem of modal differences between different modal features of media resource data, resulting in an increase in complexity in the decoder training and prediction stages.
By obtaining the target modal features of media resource data, the target mapping features corresponding to the target modal features are generated based on the target vector set, different modal features are mapped to the same vector space, and the target mapping features are decoded using a text-trained decoder to generate text descriptions.
Reduce the complexity of the decoder training and prediction stages, solve the problem of modal differences, and realize efficient text description generation of media resource data.
Smart Images

Figure CN119917679A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to computer technology, and more particularly to a method, apparatus, device and medium for generating a text description of media resource data. Background Art
[0002] By generating text descriptions for media resource data (eg, images, texts, videos, and audios, etc.), the media resource data can be better understood.
[0003] However, there are modal differences between different modal features of media resource data. If the text description task is performed through the decoder, different modal features need to be processed separately, which increases the complexity of the decoder training and prediction stages. Summary of the invention
[0004] The present disclosure provides a method, device, equipment and medium for generating text description of media resource data, which can solve the problem of modality difference and reduce the complexity of decoder training and prediction stages.
[0005] In a first aspect, an embodiment of the present disclosure provides a method for generating a text description of media resource data, comprising:
[0006] Obtain target modality features of media resource data;
[0007] Generate a target mapping feature corresponding to the target modal feature based on a target vector set, wherein the target vector set includes a set number of target vectors, and a combination of the target vectors is used to represent the target mapping feature, and the target mapping feature is a feature of the target modal feature mapped to a vector space;
[0008] The target mapping feature is decoded to generate a text description of the media resource data.
[0009] In a second aspect, the embodiments of the present disclosure further provide a device for generating a text description of media resource data, the device comprising:
[0010] A feature acquisition module, used to acquire target modality features of media resource data;
[0011] A feature mapping module, used to generate a target mapping feature corresponding to the target modal feature based on a target vector set, wherein the target vector set includes a set number of target vectors, and a combination of the target vectors is used to represent the target mapping feature, and the target mapping feature is a feature of the target modal feature mapped to the vector space;
[0012] The description generation module is used to decode the target mapping feature and generate a text description of the media resource data.
[0013] In a third aspect, an embodiment of the present disclosure further provides an electronic device, the electronic device comprising:
[0014] one or more processors;
[0015] a storage device for storing one or more programs,
[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the method for generating a text description of media resource data as described in any embodiment of the present disclosure.
[0017] In a fourth aspect, an embodiment of the present disclosure further provides a storage medium comprising computer executable instructions, which, when executed by a computer processor, are used to execute the method for generating a text description of media resource data as described in any embodiment of the present disclosure.
[0018] The disclosed embodiments provide a method, apparatus, device and medium for generating a text description of media resource data. The method obtains target modal features of the media resource data, generates target mapping features corresponding to the target modal features based on a target vector set, and maps different modal features to the same vector space to solve the modal difference problem. Then, the target mapping features are decoded to generate a text description of the media resource data. Since the decoder used for decoding is trained based on a text training data set and the text description of the media resource data is generated by the trained decoder, the complexity of the decoder training and prediction stages is reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the originals and elements are not necessarily drawn to scale.
[0020] Figure 1 A flowchart of a method for generating a text description of media resource data provided by an embodiment of the present disclosure;
[0021] Figure 2 A schematic diagram of a flow chart of an image description generation method provided by an embodiment of the present disclosure;
[0022] Figure 3 A flowchart of another method for generating a text description of media resource data provided by an embodiment of the present disclosure;
[0023] Figure 4 A flowchart of a decoder training method provided by an embodiment of the present disclosure;
[0024] Figure 5 A schematic diagram of the structure of a device for generating text description of media data provided by an embodiment of the present disclosure;
[0025] Figure 6 A schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0026] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.
[0027] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0028] The term "including" and its variations used herein are open inclusions, i.e., "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0029] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0030] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0031] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0032] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.
[0033] Figure 1A flow chart of a method for generating a text description of media resource data provided in an embodiment of the present disclosure. The method can be executed by a device for generating a text description of media resource data. The device can be implemented in the form of software and / or hardware, and optionally, by an electronic device, which can be a mobile terminal, a PC, a server, etc.
[0034] like Figure 1 As shown, the method includes:
[0035] S110: Obtain target modal features of media resource data.
[0036] The media resource data may be a description generation task to be executed to obtain data corresponding to the text description. The media resource data may include data in various forms such as images, text, video, and audio. In some embodiments, a text description may be generated for the media resource data uploaded by the user. Alternatively, a text description may be generated for the media resource data in the resource library, and a search may be performed based on the text description.
[0037] A feature of machine learning interest can be extracted from the media resource data as a target modal feature. In the disclosed embodiment, if the media resource data is an image, the image can be input into a pre-trained multimodal model, and the features of the image modal branch can be extracted as the target modal feature. If the media resource data is a video, the video can be input into a pre-trained multimodal model, and the image modal branch can be extracted as the target modal feature. Among them, the multimodal model is a neural network model trained based on sample pairs consisting of multimodal data. For example, the multimodal model can be a joint representation learning model based on images and texts, and the multimodal model is trained on sample pairs consisting of large-scale images and texts. In the disclosed embodiment, the multimodal model may include an image modal branch and a text modal branch, the image modal branch can be an image encoder, and the text modal branch can be a text encoder. The two encoders map images and texts into a representation space through contrastive learning on a large-scale image and text dataset. For images and texts with a higher degree of pairing, the cosine similarity between their features in the representation space is higher.
[0038] In some embodiments, a target image is obtained, the target image is input into a pre-trained multimodal model, and image modal features are obtained from the image modal branch of the multimodal model. For example, the multimodal model can be a joint representation learning model for images and texts, and the multimodal model includes image modal features and text modal features. By inputting the target image into the multimodal model, the image embedding and text embedding output by the model can be obtained, and the image embedding and text embedding corresponding to the target image are located in two completely independent areas in the representation space. Although the multimodal model aligns the features of the two modalities of image and text through contrastive learning, there are still very obvious modal differences between the image modality and the text modality in the representation space, which requires different modalities to be processed separately, and a unified cross-modal feature expression cannot be obtained.
[0039] S120. Generate a target mapping feature corresponding to the target modal feature based on the target vector set.
[0040] The target vector set includes a set number of target vectors. In the embodiment of the present disclosure, m sentences can be used to form a text training data set T = {T1, T2, ..., T m For example, m sentences are respectively input into the pre-trained multimodal model, and the text modality features are extracted using the text modality branch in the multimodal model. The initial vector set C = {C1, C2, ..., C n}, where C j Indicates the jth vector in the initial vector set. During the training process of the decoder, each initial vector in the initial vector set is iteratively updated, and finally the target vector set C′={C′1,C′2,……,C′ n}.
[0041] The target mapping feature is represented by a combination of the target vectors, and the target mapping feature is a feature of the target modal feature mapped to the vector space. In the disclosed embodiment, a weighted combination of target vectors in a target vector set is used to represent different modal features, so that different types of modal features can be mapped to the vector space through the target vector set.
[0042] In some embodiments, for an image or text input into a multimodal model, features corresponding to different modalities are extracted using a branch of the multimodal model, which is expressed as , where F (i) is the feature corresponding to the i-th mode, k is the total number of modes, that is, F (i) =ε (i) (·), i is the mode name.
[0043] Optionally, generating the target mapping feature corresponding to the target modal feature based on the target vector set includes: using an attention mechanism to determine the target mapping feature corresponding to the target modal feature based on the target vector in the target vector set. The attention mechanism is a method of finding the correlation between data based on the original data and then highlighting certain important features. The attention mechanism in deep learning is a method that imitates the human visual and cognitive system, which allows the neural network to focus on relevant parts when processing input data. By introducing the attention mechanism, the neural network can automatically learn and selectively focus on important information in the input. In the embodiment of the present disclosure, the feature F of any modality is calculated using the attention mechanism. (i) and each target vector C′ in the target vector set j Cosine similarity, and determine the weight of each target vector according to the cosine similarity, then use each target vector and the corresponding weight to represent the feature F of the current mode (i) .
[0044] Exemplarily, the attention mechanism is used to determine the target mapping feature corresponding to the target modal feature based on the target vector in the target vector set, including: determining the similarity between the target modal feature and the target vector in the target vector set; determining the weight of the corresponding target vector according to the similarity of the target vector; and determining the target mapping feature corresponding to the target modal feature according to the weight of the target vector and the corresponding target vector.
[0045] The similarity between the target modal feature and the target vector may be cosine similarity, Euclidean distance, etc. The weight of each target vector is obtained by normalizing the similarity of each target vector. The function used in the normalization operation may be a softmax function, etc.
[0046] In the embodiment of the present disclosure, the feature F is calculated (i) and each target vector C′ in the target vector set j Cosine similarity φ(F (i) , C′ j ), we can use the expression φ(F (i) , C′ j )=cos(F (i) , C′ j ) is expressed. Furthermore, the similarity φ(F (i) , C′ j ) After performing softmax operation, the weight a of each target vector is obtained i,j For example, the softmax operation may include the similarity φ(F (i) , C′ j) is indexed, and then weighted summation is performed, and the indexed similarity φ(F (i) , C′ j ) is normalized to obtain the weight a of each target vector i,j .
[0047] According to the weight a i,j By weighted summing and normalizing all target vectors in the target vector set, we can get the original feature F (i) Expressions in vector space.
[0048] Since the original feature F (i) It can be image features or text features, etc., and the original feature F (i) The feature expression in vector space is Therefore, regardless of image features, text features or other features, their feature expressions in the vector space are all weighted combinations of the target vectors in the target vector set. Therefore, through the target vector combination, the image modal features or text modal features can be converted into the same feature space, which can effectively solve the modal differences between different modalities in the pre-trained multimodal model.
[0049] S130: Decode the target mapping feature to generate a text description of the media resource data.
[0050] In the disclosed embodiment, decoding the target mapping feature to generate a text description of the media resource data includes: decoding the target mapping feature by a decoder to generate a text description of the media resource data.
[0051] The decoder is a neural network model trained based on a text training data set. The decoder decodes the target mapping features to obtain a text description of the media resource data. The training of the decoder does not rely on paired image and text data. Only pure text data is needed to easily complete the training task. In addition, in the inference stage, only the target mapping features of the target modal features of the media resource data in the vector space need to be input, without the need to input image and text data, which can simplify the complexity of the decoder generating text descriptions.
[0052] Figure 2 A flowchart of an image description generation method provided by an embodiment of the present disclosure is shown in FIG. Figure 2 As shown, for the input target image 1210, the image modality branch of the multimodal model 220 is used to extract the image modality features. Using the target vector set C′={C′1,C′2,……,C′ n}230 Generate F image Corresponding target mapping features Then, Input decoder 240, use decoder 240 to Decoding is performed in an autoregressive manner to obtain an image description. Since the multimodal model is composed of an image encoder and a text encoder, the target image is input into the multimodal model and the image modal features are obtained from the image modal branch, and the obtained image modal features are encoded features. Therefore, the image modal features are decoded by a decoder, and an image description can be obtained based on the decoding results, and the image description is presented in the form of a text description. The disclosed embodiment is not affected by modal differences, and only plain text data is required to easily complete the training of the decoder and generate image descriptions.
[0053] The technical solution of the disclosed embodiment obtains the target modal features of the media resource data, generates the target mapping features corresponding to the target modal features based on the target vector set, realizes mapping of different modal features to the same vector space, solves the modal difference problem, and then decodes the target mapping features through a decoder to generate a text description of the media resource data. Since the decoder is trained based on a text training data set and the trained decoder generates a text description of the media resource data, the complexity of the decoder training and prediction stages is reduced.
[0054] Figure 3 This is a flow chart of another method for generating a text description of media resource data provided by the embodiment of the present disclosure. Based on the above embodiments, this embodiment additionally defines the training method of the decoder. Figure 3 As shown, the method includes:
[0055] S310: Acquire the text training data set and the initial vector set, wherein the text training data set includes a set number of text samples.
[0056] In the embodiment of the present disclosure, a text training data set T consisting of m sentences is obtained, which is T={T1, T2, ..., T m} and obtain an initial vector set C containing n features = {C1, C2, ..., C n}. The values of m and n can be set according to actual needs, taking into account factors such as the accuracy of decoder training and resource usage.
[0057] S320. Acquire text modal features corresponding to each text sample through a pre-trained multimodal model, and generate text mapping features corresponding to the text modal features through the initial vector set, wherein the multimodal model is a neural network model trained on sample pairs composed of multimodal data.
[0058] For example, m sentences are respectively input into the pre-trained multimodal model, and the text modality feature F is extracted using the text modality branch in the multimodal model. text =ε text (T i ). Calculate F text The cosine similarity φ(F text , C j )=cos(F text , C j ), where C j represents the jth vector in the initial vector set. Then, the cosine similarity φ(F text , C j ) After performing softmax operation, we get the weight a of each vector text,j For example, the softmax operation may include the similarity φ(F text , C j ) is indexed, and then weighted summation is performed, and the indexed similarity φ(F text , C j ) is normalized to obtain the weight a of each target vector text,j .
[0059] According to the weight a text,j By weighted summing and normalizing all target vectors in the target vector set, we can get the original feature F text Mapping features in vector space, i.e. text mapping features .
[0060] S330: Using the text mapping features as input of a decoder to be trained, and supervising the training of the decoder to be trained through text samples corresponding to the text mapping features.
[0061] In the disclosed embodiment, the text sample is encoded by a multimodal model to obtain text modal features, and then a text mapping feature corresponding to the text modal features is generated using an initial vector set. Therefore, the text description obtained by decoding the text mapping feature by a decoder has a corresponding relationship with the text sample, and the decoder can be trained in a self-supervised manner based on the text sample.
[0062] Exemplarily, the text mapping features are input into a decoder to be trained; the text mapping features are decoded by the decoder to be trained to obtain a text description; an objective function value is calculated based on the text description and a text sample, and the training of the decoder to be trained is supervised based on the objective function value.
[0063] Figure 4 A flowchart of a decoder training method provided by an embodiment of the present disclosure is shown in FIG. Figure 4 As shown, for the text training data set T = {T1, T2, ..., T m For each text sample in}410, the text modality feature F is extracted using the text modality branch of the multimodal model 420 text Using the initial vector set C = {C1, C2, ..., C n}430 Generate F text Corresponding text mapping features Will As the input of the decoder to be trained 440, the text description output by the decoder to be trained 440 is obtained. According to the text description output by the decoder to be trained 440 and its corresponding text sample T i Supervise the training of the decoder to be trained 440.
[0064] At each iteration of the decoder training, the text mapping feature is input into the decoder to be trained P θ , where θ represents the parameters of the decoder. θ right Autoregressive decoding is used to obtain text description.
[0065] Get pre-built objective functions As a loss function. The objective function can be expressed as: calculate the logarithmic error between the text description corresponding to the i-th text sample and the text sample to obtain the objective function value. The objective function value represents the deviation between the text description output by the decoder and the real text sample. Then, the prediction loss of the decoder can be obtained according to the deviation, so that the parameters of the decoder are adjusted with the goal of reducing the prediction loss. For example, the parameters of the decoder to be trained are updated according to the objective function value so that the gradient of the objective function decreases. Among them, the gradient can be the rate of change of the objective function value corresponding to different decoder parameters. If the gradient decreases, the rate of change of the objective function value needs to decrease. The parameters of the decoder to be trained can be updated to decrease the gradient of the objective function. Finally, when the objective function value meets the end condition of the model training, a trained decoder is obtained.
[0066] Since the objective function The variables include the decoder parameters θ and the initial vector combination C. During the training process of the decoder, the decoder parameters are updated by gradient descent. The initial vector set can also be updated according to the objective function value, and the target vector set is determined according to the updated initialization vector set. For example, the initial vector set when the objective function value meets the end condition of model training is used as the target vector set. During the training process of the decoder, the characteristics of the target vector set are learned, so that the target vector set is used to map different modal features to a unified vector space. For example, the image and text features are uniformly mapped to the feature space of the vector set.
[0067] The technical solution of the disclosed embodiment obtains a text training data set and an initial vector set, and obtains text modal features corresponding to each text sample through a pre-trained multimodal model, generates text mapping features corresponding to the text modal features through the initial vector set, inputs the text mapping features into the decoder to be trained, and supervises the training of the decoder to be trained through text samples corresponding to the text mapping features, so that no other supervisory signals are required, and only pure text data is needed to complete the training, thereby reducing the complexity of decoder training.
[0068] Figure 5 This is a schematic diagram of the structure of a device for generating text descriptions of media data provided by an embodiment of the present disclosure. The device can be implemented in the form of software and / or hardware, and the device can execute the method for generating text descriptions of media data provided by any embodiment of the present disclosure. Figure 5 As shown, the device includes: a feature acquisition module 510, a feature mapping module 520 and a description generation module 530.
[0069] A feature acquisition module 510 is used to acquire target modality features of media resource data;
[0070] A feature mapping module 520 is used to generate a target mapping feature corresponding to the target modal feature based on a target vector set, wherein the target vector set includes a set number of target vectors, and a combination of the target vectors is used to represent the target mapping feature, and the target mapping feature is a feature of the target modal feature mapped to a vector space;
[0071] The description generation module 530 is used to decode the target mapping feature and generate a text description of the media resource data.
[0072] The technical solution provided by the disclosed embodiment obtains the target modal features of the media resource data, generates the target mapping features corresponding to the target modal features based on the target vector set, realizes mapping of different modal features to the same vector space, solves the modal difference problem, and then decodes the target mapping features through a decoder to generate a text description of the media resource data. Since the decoder is trained based on a text training data set and the trained decoder generates a text description of the media resource data, the complexity of the decoder training and prediction stages is reduced.
[0073] Optionally, the feature mapping module 520 is specifically used for:
[0074] An attention mechanism is used to determine a target mapping feature corresponding to the target modal feature based on the target vector in the target vector set.
[0075] Furthermore, the feature mapping module 520 is also specifically used for:
[0076] Determining the similarity between the target modal feature and the target vector in the target vector set;
[0077] Determining a weight of a corresponding target vector according to the similarity of the target vector;
[0078] A target mapping feature corresponding to the target modal feature is determined according to the weight of the target vector and the corresponding target vector.
[0079] Optionally, the description generating module 530 is specifically used for:
[0080] Decoding the target mapping feature by a decoder to generate a text description of the media resource data;
[0081] The device also includes a decoder training module, and the decoder training model includes:
[0082] A set acquisition unit, used to acquire the text training data set and the initial vector set, wherein the text training data set includes a set number of text samples;
[0083] A feature acquisition unit, used for acquiring text modality features corresponding to each text sample through a pre-trained multimodal model, and generating text mapping features corresponding to the text modality features through the initial vector set, wherein the multimodal model is a neural network model trained based on sample pairs composed of multimodal data;
[0084] The training unit is used to use the text mapping feature as an input of a decoder to be trained, and supervise the training of the decoder to be trained through text samples corresponding to the text mapping feature.
[0085] Furthermore, the training unit is specifically used for:
[0086] Inputting the text mapping features into a decoder to be trained;
[0087] Decoding the text mapping feature by the decoder to be trained to obtain a text description;
[0088] An objective function value is calculated according to the text description and the text sample, and the training of the decoder to be trained is supervised according to the objective function value.
[0089] Optionally, the training unit is further used for:
[0090] The parameters of the decoder to be trained are updated according to the objective function value so that the gradient of the objective function decreases.
[0091] Optionally, the device further comprises:
[0092] The set updating module is used to update the initialization vector set according to the objective function value after calculating the objective function value according to the probability, and determine the target vector set according to the updated initialization vector set.
[0093] Optionally, the feature acquisition module 510 is specifically used for:
[0094] The target image is input into a pre-trained multimodal model, and image modality features are obtained from the image modality branch of the multimodal model.
[0095] The text description generating device for media resource data provided by the embodiments of the present disclosure can execute the text description generating method for media resource data provided by any embodiment of the present disclosure, and has the functional modules and beneficial effects corresponding to the execution method.
[0096] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the protection scope of the embodiments of the present disclosure.
[0097] Figure 6 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. Figure 6 , which shows an electronic device (eg, Figure 6The terminal device in the embodiment of the present disclosure may include but is not limited to mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0098] like Figure 6 As shown, the electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An edit / output (I / O) interface 605 is also connected to the bus 604.
[0099] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 6 The electronic device 600 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.
[0100] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.
[0101] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0102] The electronic device provided in the embodiment of the present disclosure and the text description generation method of media resource data provided in the above embodiment belong to the same inventive concept. The technical details not fully described in this embodiment can be referred to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0103] The embodiment of the present disclosure provides a computer storage medium on which a computer program is stored. When the program is executed by a processor, the method for generating a text description of media resource data provided in the above embodiment is implemented.
[0104] It should be noted that the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0105] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0106] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0107] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device:
[0108] Obtain target modality features of media resource data;
[0109] Generate a target mapping feature corresponding to the target modal feature based on a target vector set, wherein the target vector set includes a set number of target vectors, and a combination of the target vectors is used to represent the target mapping feature, and the target mapping feature is a feature of the target modal feature mapped to a vector space;
[0110] The target mapping feature is decoded to generate a text description of the media resource data.
[0111] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or a combination thereof, including, but not limited to, object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0112] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0113] The units involved in the embodiments described in the present disclosure may be implemented by software or hardware, wherein the name of a unit does not, in some cases, limit the unit itself.
[0114] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0115] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0116] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present disclosure (but not limited to) by each other to form a technical solution.
[0117] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0118] Although the subject matter has been described in language specific to structural features and / or methodological logical actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely example forms of implementing the claims.
Claims
1. A method for generating a text description of media resource data, characterized in that: include: Obtain target modality features of media resource data; Generate a target mapping feature corresponding to the target modal feature based on a target vector set, wherein the target vector set includes a set number of target vectors, and a combination of the target vectors is used to represent the target mapping feature, and the target mapping feature is a feature of the target modal feature mapped to a vector space; The target mapping feature is decoded to generate a text description of the media resource data.
2. The method according to claim 1, characterized in that: The generating the target mapping feature corresponding to the target modal feature based on the target vector set includes: An attention mechanism is used to determine a target mapping feature corresponding to the target modality feature based on the target vector in the target vector set.
3. The method according to claim 2, characterized in that The using of the attention mechanism to determine the target mapping feature corresponding to the target modality feature based on the target vector in the target vector set includes: Determining the similarity between the target modal feature and the target vector in the target vector set; Determining a weight of a corresponding target vector according to the similarity of the target vector; A target mapping feature corresponding to the target modal feature is determined according to the weight of the target vector and the corresponding target vector.
4. The method according to claim 1, characterized in that Decoding the target mapping feature to generate a text description of the media resource data also includes: Decoding the target mapping feature by a decoder to generate a text description of the media resource data; The decoder is trained as follows: Acquire the text training data set and the initial vector set, wherein the text training data set includes a set number of text samples; Acquire text modality features corresponding to each text sample through a pre-trained multimodal model, and generate text mapping features corresponding to the text modality features through the initial vector set, wherein the multimodal model is a neural network model trained based on sample pairs composed of multimodal data; The text mapping features are used as input of a decoder to be trained, and the training of the decoder to be trained is supervised by text samples corresponding to the text mapping features.
5. The method according to claim 4, characterized in that The step of using the text mapping feature as an input of a decoder to be trained, and supervising the training of the decoder to be trained through text samples corresponding to the text mapping feature, comprises: Inputting the text mapping features into a decoder to be trained; Decoding the text mapping feature by the decoder to be trained to obtain a text description; An objective function value is calculated according to the text description and the text sample, and the training of the decoder to be trained is supervised according to the objective function value.
6. The method according to claim 5, characterized in that The supervising the training of the decoder to be trained according to the objective function value comprises: The parameters of the decoder to be trained are updated according to the objective function value so that the gradient of the objective function decreases.
7. The method according to claim 5, characterized in that After calculating the objective function value according to the probability, it also includes: The initialization vector set is updated according to the objective function value, and the target vector set is determined according to the updated initialization vector set.
8. The method according to claim 4, characterized in that The step of acquiring target modality features of media resource data includes: The target image is input into a pre-trained multimodal model, and image modality features are obtained from the image modality branch of the multimodal model.
9. A device for generating text description of media resource data, characterized in that: include: A feature acquisition module, used to acquire target modality features of media resource data; A feature mapping module, used to generate a target mapping feature corresponding to the target modal feature based on a target vector set, wherein the target vector set includes a set number of target vectors, and a combination of the target vectors is used to represent the target mapping feature, and the target mapping feature is a feature of the target modal feature mapped to the vector space; The description generation module is used to decode the target mapping feature and generate a text description of the media resource data.
10. An electronic device, characterized in that: The electronic device comprises: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method for generating a text description of media resource data as described in any one of claims 1-8.
11. A storage medium containing computer executable instructions, characterized in that: The computer executable instructions are used to execute the text description generation method of media resource data as described in any one of claims 1-8 when executed by a computer processor.