Method, device, storage medium and processor for processing multimedia information
By identifying and fusing information from different modalities of multimedia information in videos, and using a self-attention mechanism to generate recommended content, the problem of low recognition efficiency caused by the complexity of multimodal information is solved, and more efficient recommended content recognition is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-02
- Publication Date
- 2026-03-27
AI Technical Summary
In existing technologies, the complexity of multimodal information leads to low efficiency in identifying key recommended content in videos and makes it impossible to effectively fuse multimodal information.
By acquiring multimedia information from videos, identifying information about target objects in different modalities, and employing a self-attention mechanism for multimodal fusion, recommended content is generated to represent the target objects.
It improves the recognition efficiency of key recommended content in videos and solves the problem of low recognition efficiency caused by the complexity of multimodal information.
Smart Images

Figure CN114443938B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of information processing, and in particular, to a multimedia information processing method and device, a storage medium and a processor. BACKGROUND
[0002] At present, with the rise of live streaming of goods on some transaction platforms in e-commerce platforms, video technology is also increasingly widely used in modern society. As a new interactive medium, video not only contains rich image information, but also presents multi-modal information such as text and audio. Generally, multi-modal information can be fully utilized to obtain better performance than single image, but due to the complexity of multi-modal information, it is difficult to effectively fuse multi-modal information, thereby causing the technical problem of low efficiency in identifying key recommended content in the video.
[0003] At present, there is no effective solution to the problem of low efficiency in identifying key recommended content in the video. SUMMARY
[0004] The embodiments of the present application provide a multimedia information processing method and device, a storage medium and a processor to at least solve the technical problem of low efficiency in identifying key recommended content in the video.
[0005] According to one aspect of the embodiments of the present application, a multimedia information processing method is provided, comprising: playing a video and obtaining multimedia information in the video, wherein the multimedia information includes image information and audio information; identifying information of a target object in different modalities from the multimedia information, wherein the target object is a product recommended in the video; and performing multi-modal fusion on the information of the target object in different modalities to generate recommended content for representing the target object.
[0006] According to another aspect of the embodiments of the present application, a multimedia information processing method is also provided, comprising: entering multimedia information in a played video in an entry interface of an operation interface, wherein the multimedia information includes image information and audio information; sensing a recommended content generation instruction in the operation interface and identifying information of a target object in different modalities from the multimedia information, wherein the target object is a product recommended in the video; and displaying recommended content for representing the target object on the operation interface, wherein the recommended content for representing the target object is generated by performing multi-modal fusion on the information of the target object in different modalities.
[0007] According to another aspect of the embodiments of the present application, a method for processing multimedia information is also provided, including: displaying a played video and multimedia information in the video in an operation interface, wherein the multimedia information includes image information and audio information; sensing a recommendation content generation instruction in the operation interface; displaying information of a target object in different modalities on the operation interface in response to the content generation instruction, wherein the target object is a recommended product in the video, and the information of the target object in different modalities is identified from the multimedia information; and displaying recommendation content for representing the target object on the operation interface, wherein the recommendation content for representing the target object is generated by performing multi-modal fusion on the information of the target object in different modalities.
[0008] According to another aspect of the embodiments of the present application, a method for processing multimedia information is also provided, including: uploading multimedia information in a played video by a front-end client, wherein the multimedia information includes image information and audio information; transmitting the multimedia information to a back-end server by the front-end client; receiving information of a target object in different modalities returned by the back-end server from the multimedia information by the front-end client, wherein the target object is a recommended product in the video; and performing multi-modal fusion on the information of the target object in different modalities by the front-end client to generate recommendation content for representing the target object.
[0009] According to another aspect of the embodiments of the present application, a method for processing multimedia information is also provided, including: receiving a recommendation content generation request; obtaining multimedia information of a video in the recommendation content generation request, wherein the multimedia information includes image information and audio information; identifying information of a target object in different modalities from the multimedia information, wherein the target object is a recommended product in the video; performing multi-modal fusion on the information of the target object in different modalities to generate recommendation content for representing the target object; and outputting the recommendation content for representing the target object.
[0010] According to another aspect of the embodiments of the present application, a device for processing multimedia information is also provided, including: a playing unit configured to play a video and obtain multimedia information in the video, wherein the multimedia information includes image information and audio information; an identifying unit configured to identify information of a target object in different modalities from the multimedia information, wherein the target object is a recommended product in the video; and a fusion unit configured to perform multi-modal fusion on the information of the target object in different modalities to generate recommendation content for representing the target object.
[0011] According to another aspect of the embodiments of the present application, a live broadcast all-in-one machine is also provided, comprising: a player configured to play a video and acquire multimedia information in the video, wherein the multimedia information comprises image information and audio information; an identifier connected to the player and configured to identify information of a target object in different modalities from the multimedia information, wherein the target object is a product recommended in the video; and a generator connected to the identifier and configured to perform multi-modal fusion on the information of the target object in different modalities to generate recommendation content for representing the target object.
[0012] According to another aspect of the embodiments of the present application, a computer readable storage medium is also provided. The computer readable storage medium can include a stored program, wherein the program, when executed by a processor, controls a device where the computer readable storage medium is located to perform the multimedia information processing method according to the embodiments of the present application.
[0013] According to another aspect of the embodiments of the present application, a processor is also provided. The processor is configured to execute a program, wherein the program, when executed, performs the multimedia information processing method according to the embodiments of the present application.
[0014] According to another aspect of the embodiments of the present application, a multimedia information processing system is also provided. The multimedia information processing system can include: a processor; and a memory connected to the processor and configured to provide the processor with instructions for processing the following processing steps: playing a video and acquiring multimedia information in the video, wherein the multimedia information comprises image information and audio information; identifying information of a target object in different modalities from the multimedia information, wherein the target object is a product recommended in the video; and performing multi-modal fusion on the information of the target object in different modalities to generate recommendation content for representing the target object.
[0015] In the embodiment of the present application, the video is played, and multimedia information in the video is acquired, wherein the multimedia information includes image information and audio information; information of a target object in different modalities is identified from the multimedia information, wherein the target object is a recommended product in the video; and the information of the target object in different modalities is subjected to multi-modal fusion to generate recommendation content for representing the target object. In the related art, due to the low abstract degree of the image, richer information can be exhibited, but the image is also more difficult to be understood by the computer, therefore, auxiliary information in additional modalities such as audio and text is usually subjected to vectorization in a manner similar to word2vector, and a recurrent neural network is used to acquire the features of the auxiliary modalities, and in the modal fusion, information between the interactive modalities is usually fused in a manner of simultaneous fusion / training, so as to acquire a more robust video representation. This method only needs to map the image content into text and then structure and organize the text when converting the image into the text, but due to the complexity of the multi-modal information, the multi-modal information cannot be effectively fused, and in the present application, the multimedia information in the video is acquired to identify the information of the target object in different modalities in the multimedia information, so that the information of the target object in different modalities can be subjected to multi-modal fusion, so that the recommendation content for representing the target object can be generated, thereby avoiding the problem that the complexity of the multi-modal information is high, and the efficiency of identifying the key recommendation content in the video is low, and further solving the technical problem that the efficiency of identifying the key recommendation content in the video is low, and achieving the technical effect of improving the efficiency of identifying the key recommendation content in the video. BRIEF DESCRIPTION OF DRAWINGS
[0016] The accompanying drawings, which are included to provide a further understanding of the application and are incorporated in and constitute a part of this application, illustrate embodiments of the application and together with the description serve to explain the application. In the drawings:
[0017] Figure 1 is a hardware structure block diagram of a computer terminal (or mobile device) for a multimedia information processing method according to an embodiment of the present application;
[0018] Figure 2 is a flowchart of a multimedia information processing method according to an embodiment of the present application;
[0019] Figure 3 is a flowchart of another multimedia information processing method according to an embodiment of the present application;
[0020] Figure 4 is a schematic diagram of a query intention (Query) extraction method according to an embodiment of the present application;
[0021] Figure 5is a scene schematic diagram of processing of multimedia information according to an embodiment of the present application;
[0022] Figure 6 is a schematic diagram of a processing device of multimedia information according to an embodiment of the present application; and
[0023] Figure 7 is a structural block diagram of a mobile terminal according to an embodiment of the present application. DETAILED DESCRIPTION
[0024] In order to make the personnel in the art better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should belong to the scope of protection of the present application.
[0025] It should be noted that the terms "first", "second" and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0026] First, some of the nouns or terms appearing in the description of the embodiments of the present application are applicable to the following explanations:
[0027] Query: used to find a specific file, website, record or a series of records in a database, a message sent by a search engine or a database, used in the present application to understand the goods explained by the host in live broadcast;
[0028] Convolutional Neural Network (CNN): a kind of feedforward neural network, the artificial neuron of which can respond to a part of the surrounding units in the coverage range, which is a deep learning method for image recognition;
[0029] Automatic Speech Recognition (ASR): an automatic speech-to-text system that converts human speech into text.
[0030] Bidirectional Encoder Representations from Transformers (BERT): a deep model for learning semantic features of text.
[0031] Embodiment 1
[0032] According to the embodiments of the present application, an embodiment of a processing method of multimedia information is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0033] The method embodiment provided by the embodiment one of the present application can be executed in a mobile terminal, a computer terminal or similar computing device. Figure 1 is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a processing method of multimedia information according to the embodiments of the present application. As shown in Figure 1 , the computer terminal 10 (or mobile device 10) can include one or more processors 102 (the processor 102 can include but not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication function. In addition, it can also include a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports of the I / O interface), a network interface, a power supply and / or a camera. Those skilled in the art can understand that Figure 1 The structure shown is only schematic, which does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 can also include more or less components than those shown in Figure 1 , or have a different configuration from Figure 1 .
[0034] It should be noted that the one or more processors 102 and / or other data processing circuitry described above can be generally referred to herein as "data processing circuitry." The data processing circuitry can be embodied in whole or in part as software, hardware, firmware, or any combination thereof. In addition, the data processing circuitry can be a single standalone processing module, or incorporated in whole or in part within any of the other elements of the computer terminal 10 (or mobile device). As referred to in embodiments of the present application, the data processing circuitry functions as a processor to control, for example, selection of variable resistance terminal paths connected to the interface.
[0035] The memory 104 can be used to store software programs of application software and modules, such as program instructions / data storage means corresponding to the multimedia information processing method of embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, i.e., implements the multimedia information processing method described above. The memory 104 can include a high-speed random access memory, and can further include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 can further include a memory disposed remotely with respect to the processor 102, which can be connected to the computer terminal 10 through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0036] The transmission device 106 is used to receive or send data via a network. Specific examples of the network can include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network adapter (NIC) that can be connected to other network devices through a base station so as to be able to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module used to communicate with the Internet in a wireless manner.
[0037] The display can be, for example, a touch screen type liquid crystal display (LCD) that can enable a user to interact with the user interface of the computer terminal 10 (or mobile device).
[0038] It should be noted that in some alternative embodiments, the above Figure 1 The computer device (or mobile device) shown can include hardware elements (including circuitry), software elements (including computer code stored on a computer readable medium), or a combination of both hardware and software elements. It should be noted that in some embodiments, the functions of the computer device (or mobile device) described above can be combined in one computer device (or mobile device), or divided into separate computer devices (or mobile devices). Figure 1is merely one instance of a particular, concrete example, and is intended to show the types of components that can be present in the above-described computer device (or mobile device).
[0039] In Figure 1 The present application provides a multimedia information processing method as shown in Figure 2 The present application provides a multimedia information processing method as shown in Figure 1 The present application provides a multimedia information processing method as shown in
[0040] Figure 2 is a flowchart of a multimedia information processing method according to an embodiment of the present application. As shown in Figure 2 The method can include the following steps:
[0041] Step S202: Play a video and obtain multimedia information in the video.
[0042] In the technical solution provided by the above step S202 of the present application, the multimedia information in the video can be obtained by playing the video, and the obtained multimedia information can include image information and audio information.
[0043] Optionally, the video in this embodiment can be a live video, and the multimedia information can include but is not limited to live image information and voice information.
[0044] Step S204: Identify information of a target object in different modalities from the multimedia information.
[0045] In the technical solution provided by the above step S204 of the present application, after obtaining the multimedia information in the video, the information of the target object in different modalities can be identified from the multimedia information. In this embodiment, the target object can be a product recommended in the video, and the information of the target object in different modalities can include information such as the type, length, size, color, structure, and function of the product.
[0046] In this embodiment, the information of the target object in different modalities can include information of the target object in modalities such as a voice modality, a video modality, a picture modality, and a text modality. For example, the information of the target object in the voice modality can include information of the host describing or recommending the product through voice, the information of the target object in the video modality can include information of the product displayed through the video, the information of the target object in the picture modality can include information of the product displayed through the picture, and the information of the target object in the text modality can include information of the host describing or recommending the product through text.
[0047] In step S206, the information of the target object in different modalities is fused to generate recommended content for representing the target object.
[0048] In the technical solution provided in step S206, after obtaining the information of the target object in different modalities, the information of the target object in different modalities is fused, so that the recommended content for representing the target object is generated, that is, the content for recommending the target object.
[0049] In this embodiment, the fusion method of the self-attention mechanism can be used to fuse the information of the target object in different modalities. Optionally, the features of the information of the target object in different modalities can be normalized, and then the normalized features are spliced and fused.
[0050] In the related art, when converting an image into text, only the image content is mapped into text and then structured and arranged. However, due to the high complexity of multi-modal information, the multi-modal information cannot be effectively fused.
[0051] However, by using steps S202 to S206, the video is played, and the multi-media information in the video is obtained, where the multi-media information includes image information and audio information. The information of the target object in different modalities is identified from the multi-media information, where the target object is a product recommended in the video. The information of the target object in different modalities is fused to generate recommended content for representing the target object. That is, by obtaining the multi-media information in the video, the information of the target object in different modalities in the multi-media information can be identified, so that the information of the target object in different modalities can be fused, and the recommended content for representing the target object can be generated. Therefore, the problem of low efficiency in identifying the key recommended content in the video due to the high complexity of multi-modal information is avoided, and the technical problem of low efficiency in identifying the key recommended content in the video is solved, and the technical effect of improving the efficiency of identifying the key recommended content in the video is achieved.
[0052] The above method of this embodiment will be further described below.
[0053] As an optional implementation, in step S204, identifying the information of the target object in different modalities from the multi-media information includes: performing image sampling on the played video to obtain image information, where the image information includes a video frame sequence composed of image frames; and using a detector to detect the video frame sequence to obtain visual feature information of the video on a visual track in a playing process, where the visual feature information includes image dimension information of the target object in different dimensions.
[0054] In this embodiment, when identifying the multimedia information, the played video is image-sampled, which can be equidistant image-sampling of the played video. Since the video is a video in a playing state, the obtained image information can be a video frame sequence composed of image frames with time sequence. Then, a detector can be used to detect the video frame sequence to obtain visual feature information of the visual track of the video in the playing process. The visual feature information can include image dimension information of the target object in different dimensions.
[0055] Optionally, when image-sampling the played video, a time interval can be preset, and the played video can be image-sampled once every time interval.
[0056] Optionally, the detector in this embodiment can include a RetinaNet detector, a YoloV3 detector, etc.
[0057] As an optional implementation, the detector is used to detect the video frame sequence to obtain visual feature information of the visual track of the video in the playing process, including: detecting the video frame sequence to obtain bounding box information of at least one bounding box in the video; and identifying image dimension information of a target object played in the video based on the bounding box information of the bounding box.
[0058] In this embodiment, after obtaining the video frame sequence of the video, a detector can be used to detect the obtained video frame sequence to obtain bounding box information of at least one bounding box in the video. Then, according to the bounding box information, the image dimension information of the target object in the playing process of the video is identified.
[0059] In the above embodiment, the image dimension information can include at least one of the following: trajectory information of the target object displayed in the video frame sequence, position coordinates of the target object in each video frame, and feature information of the target object.
[0060] Optionally, although the bounding box in this embodiment can effectively identify the image dimension information in the video, it cannot effectively solve the problems of image inseparability, image partial separability, and non-key region interference in the image. Therefore, the present application can also track multiple commodities (Min-cost Flow, referred to as MCF) to associate multiple commodities to obtain information such as bounding boxes, track features, and human key points at different time sequences.
[0061] As an optional implementation, the step S204 comprises: performing audio sampling on the played video to obtain audio information, wherein the audio information comprises an audio frame sequence composed of audio frames; and converting the audio information into text information, wherein the text information describes the textual feature information of the target object played in the video.
[0062] In this embodiment, when the multimedia information is identified, the played video can be audio sampled. Since in the live video, the audio information can effectively describe the product being explained by the anchor in the video, and can reflect the saliency region that is not easy to distinguish in the image, and can judge the effectiveness of the current video frame, eliminate the video frame with invalid content, and can also describe the product features that are not easy to distinguish in the image, such as the style of the product, thereby helping to extract the features of the image, and then the effective audio information can be obtained, which can include an audio frame sequence composed of audio frames. Then the obtained audio information is converted into text information, so as to describe the target object played in the video frame by the textual feature information.
[0063] Optionally, in this embodiment, the audio information can be converted into text information by using automatic speech recognition technology.
[0064] Optionally, the text information in this embodiment can include the style, category, color and other information of the product.
[0065] Optionally, in this embodiment, the text splicing can be performed in advance based on the textual feature information, and the spliced features are taken as the input of the pre-trained BERT model.
[0066] As an optional implementation, after the audio information is converted into text information, the method further comprises: performing word segmentation processing on the text information to obtain at least one keyword used to describe the target object; and determining the content and the part of speech of the keyword.
[0067] In this embodiment, after the audio information is converted into text information, the text information can be subjected to word segmentation processing, so that at least one keyword used to describe the target object can be obtained, and then the content and the part of speech of the obtained keyword can be determined, so as to describe the target object according to the keyword, thereby avoiding the problem that after the audio information is converted into text information, the content of the obtained text information is too much and the information is too complex, resulting in difficulty in reading the semantic information of the target object, and thus the purpose of improving the efficiency of reading the information of the target object in the video can be achieved.
[0068] As an optional implementation, in step S206, the information of the target object in different modalities is fused to generate recommended content for representing the target object, including: generating a feature set in an image modality based on visual feature information of the target object; generating a feature set in a text modality based on textual feature information of the target object; and generating recommended content of the target object based on the feature set in the image modality and the feature set in the text modality.
[0069] In this embodiment, for the visual feature information of the target object on the visual track, a convolutional neural network algorithm can be used to obtain the feature set in the image modality; for the textual feature information of the target object, a BERT encoding can be used to obtain the feature set in the text modality, and then the feature set in the image modality and the feature set in the text modality are jointly input into a self-attention mechanism module to obtain joint feature embedding, so that the final decoding result is obtained through the response of the textual feature information to the visual feature information on the visual track, and then the recommended content of the target object is generated.
[0070] In the above embodiment, the feature set in the image modality and the feature set in the text modality are jointly input into the self-attention mechanism module, and there are two input parts, one is a feature map in the image modality, and the other is a feature vector in the text modality. In this embodiment, the feature vector in the text modality is used to assist in learning the weighting matrix for the feature map in the spatial domain. The calculation method of the learning is as follows: first, the similarity or correlation between the query intention and each keyword is calculated (here, the similarity or correlation between the query intention and each keyword can be obtained by point multiplication of the query intention and each keyword), to obtain the weight coefficient of the value corresponding to each keyword, and then the weighted sum of the value is obtained, that is, the final attention value is obtained. Therefore, the essence of the attention mechanism is to perform weighted sum on the values of the elements in the source modality feature, and the query intention and the keyword are used to calculate the weight coefficient of the corresponding value, and each keyword in each sentence needs to perform attention calculation with all the words in the sentence, so as to realize the learning of the word dependency relationship in the sentence and capture the internal structure of the sentence.
[0071] As an optional implementation, the visual feature information is processed by a neural network model to obtain the feature set in the image modality.
[0072] In this embodiment, since the image data has higher dimension and more noise, and the abstraction degree of the image is low, it is not easy for the computer to understand, so the neural network model can be used to process the visual feature information to obtain the feature set in the image modality, so as to improve the efficiency of processing the visual feature information of the image.
[0073] As an optional implementation, the visual feature information is encoded by BERT encoding to obtain a feature set in a text mode.
[0074] In this embodiment, since the text data has less noise and the text is structured and has certain grammar rules, the visual feature information can be encoded by BERT encoding to obtain a feature set in a text mode, so as to improve the efficiency of processing the text feature information.
[0075] As an optional implementation, the feature set in the image mode and the feature set in the text mode are fused to generate the recommended content of the target object.
[0076] In this embodiment, after obtaining the feature set in the image mode and the feature set in the text mode, the two feature sets can be fused to generate the recommended content of the target object.
[0077] Optionally, the embodiment can select a fusion method based on a self-attention mechanism to fuse the feature set in the image mode and the feature set in the text mode.
[0078] As an optional implementation, the video is a process in which a host explains a target object, and in the fusion process, whether the feature set in the image mode and the feature set in the text mode have a corresponding relationship is judged to determine whether the target object played is the object explained by the host.
[0079] In this embodiment, whether the visual track of the product in the live video is the product explained by the host can be determined by the fusion result obtained by the fusion method based on the self-attention mechanism, that is, whether the target object played is the object explained by the host can be determined by judging whether the feature set in the image mode and the feature set in the text mode have a corresponding relationship. For example, this method is similar to simplifying a visual question and answer problem into a binary classification model with a predictable answer range, and the training method can be that two data are obtained, the label of the target product explained by the host can be 1, and the labels of the remaining background products or products not explained by the host are 0.
[0080] The embodiment of the application also provides another processing method of multimedia information from the perspective of human-computer interaction. The method can include: entering multimedia information in a video played in an entry interface of an operation interface, wherein the multimedia information includes image information and audio information; sensing a recommended content generation instruction in the operation interface, identifying information of a target object in different modes from the multimedia information, wherein the target object is a product recommended in the video; and displaying recommended content for representing the target object on the operation interface, wherein the recommended content of the target object is generated by multi-modal fusion of the information of the target object in different modes.
[0081] In this embodiment, the multimedia information in the played video can be entered in the entry interface of the operation interface, which can include image information and audio information, and when the operation interface receives a recommended content generation instruction, the multimedia information in the entered played video can be identified according to the indication of the instruction, so as to obtain the information of the target object in different modalities, and then the information of the target object in different modalities is multi-modal fusion, thereby generating the recommended content of the target object, and displaying the recommended content on the operation interface.
[0082] In the above embodiment, the information of the target object in different modalities can include information of the target object in voice modalities, video modalities, picture modalities, text modalities and the like. For example, the information of the target object in voice modalities can include information of the host describing or recommending the product through voice, the information of the target object in video modalities can include information of the product displayed through video, the information of the target object in picture modalities can include information of the product displayed through picture, and the information of the target object in text modalities can include information of the product described or recommended through text.
[0083] Optionally, the embodiment can use a self-attention mechanism fusion method to realize multi-modal fusion of the information of the target object in different modalities. For example, the features of the information of the target object in different modalities can be normalized, and then the normalized features are spliced and fused.
[0084] Optionally, the target object in the embodiment can be a product recommended in the played video, and the information of the target object in different modalities can include information of the type, length, size, color, structure, function and the like of the product.
[0085] The embodiment of the application also provides another multimedia information processing method from the perspective of human-computer interaction. The method can include: displaying a played video in an operation interface, and displaying multimedia information in the video, wherein the multimedia information includes image information and audio information; sensing a recommended content generation instruction in the operation interface; in response to the content generation instruction, displaying information of a target object in different modalities on the operation interface, wherein the target object is a product recommended in the video, and the information of the target object in different modalities is identified from the multimedia information; displaying recommended content for representing the target object on the operation interface, wherein the recommended content of the target object is generated by multi-modal fusion of the information of the target object in different modalities.
[0086] In this embodiment, the played video can be displayed through the operation interface, and multimedia information such as image information and audio information in the video can be displayed. When the operation interface receives a recommended content generation instruction, the instruction can be responded to, and information of the target object in different modalities and recommended content for representing the target object can be displayed on the operation interface.
[0087] In the above embodiment, after the multimedia information in the video is displayed, information of the target object in different modalities can be identified from the multimedia information, and the information of the target object in different modalities can be multi-modal fusion, so as to generate the recommended content for representing the target object.
[0088] In the above embodiment, the information of the target object in different modalities can include information of the target object in a voice modality, a video modality, a picture modality, a text modality, and the like.
[0089] Optionally, the embodiment can use a fusion method of a self-attention mechanism to realize multi-modal fusion of the information of the target object in different modalities.
[0090] Optionally, the target object in the embodiment can be a product recommended in the played video, and the information of the target object in different modalities can include information of a type, a length, a size, a color, a structure, a function, and the like of the product.
[0091] The embodiment of the application also provides another multimedia information processing method from the user's perspective. The method can include: uploading multimedia information in a played video by a front-end client, wherein the multimedia information includes image information and audio information; transmitting the multimedia information to a back-end server by the front-end client; receiving information of a target object in different modalities identified from the multimedia information returned by the back-end server, wherein the target object is a product recommended in the video; and multi-modal fusion of the information of the target object in different modalities by the front-end client, to generate recommended content for representing the target object.
[0092] In this embodiment, the multimedia information in the played video can be uploaded by the front-end client. The multimedia information can be image information or audio information. After the front-end client uploads the multimedia information in the played video, the multimedia information can be transmitted to the back-end server. After information of the target object in different modalities is identified from the multimedia information, the back-end server can return the identified information of the target object in different modalities to the front-end client. After the front-end client receives the information, the information of the target object in different modalities can be multi-modal fusion, so as to generate the recommended content for representing the target object.
[0093] In the above embodiment, the information of the target object in different modalities can include information of the target object in a voice modality, a video modality, a picture modality, a text modality, etc.
[0094] Optionally, the embodiment can adopt a fusion method of a self-attention mechanism to realize multi-modal fusion of the information of the target object in different modalities.
[0095] Optionally, the target object in the embodiment can be a product recommended in a played video, and the information of the target object in different modalities can include information of a type, a length, a size, a color, a structure, a function, etc. of the product.
[0096] The embodiment of the application further provides another multimedia information processing method. The method can include: receiving a recommendation content generation request; obtaining multimedia information of a video in the recommendation content generation request, wherein the multimedia information includes image information and audio information; identifying information of a target object in different modalities from the multimedia information, wherein the target object is a product recommended in the video; performing multi-modal fusion on the information of the target object in different modalities to generate recommendation content for representing the target object; and outputting the recommendation content of the target object.
[0097] In the embodiment, in order to better process the multimedia information, the server can receive a recommendation content generation request, and then obtain multimedia information of a video in the recommendation content generation request. The multimedia information can be image information or audio information. After the server obtains the multimedia information of the video, the information of a target object in different modalities can be identified from the multimedia information. Then, multi-modal fusion is performed on the identified information of the target object in different modalities, so as to generate recommendation content for representing the target object, and the server outputs the recommendation content of the target object.
[0098] In the above embodiment, the information of the target object in different modalities can include information of the target object in a voice modality, a video modality, a picture modality, a text modality, etc.
[0099] Optionally, the embodiment can adopt a fusion method of a self-attention mechanism to realize multi-modal fusion of the information of the target object in different modalities.
[0100] Optionally, the target object in the embodiment can be a product recommended in a played video, and the information of the target object in different modalities can include information of a type, a length, a size, a color, a structure, a function, etc. of the product.
[0101] The embodiment of the present application also provides a live broadcast integrated machine, comprising: a player configured to play a video and acquire multimedia information in the video, wherein the multimedia information comprises image information and audio information; an identifier connected to the player and configured to identify information of a target object in different modalities from the multimedia information, wherein the target object is a recommended product in the video; and a generator connected to the identifier and configured to perform multi-modal fusion on the information of the target object in different modalities to generate recommendation content for representing the target object.
[0102] In the embodiment, the live broadcast integrated machine can comprise the player, the identifier and the generator, wherein the player can be configured to play the video and acquire multimedia information in the played video, the multimedia information can be image information or audio information; the identifier can be connected to the player and configured to identify information of a target object in different modalities from the multimedia information acquired by the player; and the generator can be connected to the identifier and configured to perform multi-modal fusion on the information of the target object in different modalities, thereby generating recommendation content for representing the target object.
[0103] In the above embodiment, the information of the target object in different modalities can comprise information of the target object in a voice modality, a video modality, a picture modality, a text modality and the like.
[0104] Optionally, the embodiment can use a fusion method of a self-attention mechanism to perform multi-modal fusion on the information of the target object in different modalities.
[0105] Optionally, the target object in the embodiment can be a recommended product in the played video, and the information of the target object in different modalities can comprise information of a type, a length, a size, a color, a structure and a function of the product.
[0106] In the related art, since the abstraction degree of an image is low, the image can exhibit more abundant information, but the image is also more difficult to be understood by a computer, therefore, auxiliary information in additional modalities such as audio and text is usually vectorized in a manner similar to a word vector (word2vector), and a recurrent neural network is used to acquire features of the auxiliary modalities, and in modal fusion, information between the modalities is usually interacted during post-fusion / training, thereby acquiring a more robust video representation. This method only needs to map image content into text and then structure and organize the text when converting the image into the text, but due to the high complexity of multi-modal information, the multi-modal information cannot be effectively fused.
[0107] The embodiment can obtain multimedia information in the video, identify information of the target object in the multimedia information in different modalities, and perform multi-modal fusion on the information of the target object in different modalities, so as to generate recommended content for representing the target object, thereby avoiding the problem of low efficiency in identifying key recommended content in the video due to high complexity of multi-modal information, and solving the technical problem of low efficiency in identifying key recommended content in the video, and achieving the technical effect of improving the efficiency of identifying key recommended content in the video.
[0108] Embodiment 2
[0109] The technical solutions of the present application will be further introduced below by the method of learning query intention understanding model through two channels of vision and text.
[0110] Figure 3 is a flowchart of another method for processing multimedia information according to an embodiment of the present application. As shown in Figure 3 , the method can include the following steps:
[0111] Step S3102, acquiring live visual information.
[0112] In the technical solution provided in the above step S3102 of the present application, the live visual information of the target object in the live video can be acquired during the process of the anchor person performing live broadcast.
[0113] Optionally, the target object in the embodiment can include a commodity or a product in the live video.
[0114] Step S3104, detecting the video frame.
[0115] In the technical solution provided in the above step S3104 of the present application, after the live visual information is acquired, the live video can be sampled at equal intervals to obtain a video frame sequence with time sequence, and then a detector can be used to detect the video frame.
[0116] Optionally, the detector in the embodiment can include a RetinaNet detector, a YoloV3 detector, etc.
[0117] Step S3106, acquiring trajectory information, region of interest coordinates, and human body key points of the corresponding commodity.
[0118] In the technical solution provided in the above step S3106 of the present application, after the video frame of the live video is detected, the trajectory information, the region of interest coordinates, and the human body key points of the corresponding commodity in the live video can be acquired.
[0119] In this embodiment, the trajectory information of the commodity can include visual feature information of the commodity on the visual trajectory in the live broadcast process, and the visual feature information can include image dimension information of the commodity in different dimensions.
[0120] Optionally, the embodiment can associate multiple commodities through a multi-commodity tracking method to obtain trajectory information, coordinates of a region of interest, and human body key points of the corresponding commodity.
[0121] Step S3108, feature splicing.
[0122] In the technical solution provided by the above step S3108 of the application, after obtaining the trajectory information, coordinates of a region of interest, and human body key points of the commodity, these features can be spliced, and the spliced features can be used as an input of a query intent classifier of a self-attention mechanism model.
[0123] Step S3202, obtaining anchor speech information.
[0124] In the technical solution provided by the above step S3202 of the application, the anchor speech information can be obtained during the live broadcast of the anchor. The speech information can effectively describe the product being explained by the anchor in the live broadcast video, can reflect the saliency region that is not easy to distinguish in the image, can judge the effectiveness of the current video frame, eliminate the video frame with invalid content, and can describe the commodity features that are not easy to distinguish on the image, such as the style of the product.
[0125] Step S3204, performing speech recognition on the audio in the speech information to convert the audio into text.
[0126] In the technical solution provided by the above step S3204 of the application, after obtaining the anchor speech information, the audio in the obtained speech information can be converted into text information, so as to describe the commodity in the live broadcast video through the text feature information.
[0127] Optionally, in this embodiment, the audio information can be converted into text information through an automatic speech recognition technology.
[0128] Optionally, the text information in this embodiment can include information such as the style, category, and color of the product.
[0129] Step S3206, performing word segmentation and part-of-speech extraction on the text to obtain keywords of the commodity.
[0130] In the technical solution provided in the above step S3206 of the present application, after the audio information is converted into text information, the text information can be segmented, and the parts of speech in the text information can be extracted, so that the keywords of the commodity can be obtained, and the keywords are used to describe the commodity information.
[0131] In the technical solution provided in the above step S3208 of the present application, after the keywords of the commodity are obtained, the BERT features of the keywords can be extracted, so that the semantic information of the keywords can be obtained.
[0132] In the technical solution provided in the above step S3208 of the present application, after the keywords of the commodity are obtained, the BERT features of the keywords can be extracted, so that the semantic information of the keywords can be obtained.
[0133] In the technical solution provided in the above step S3300 of the present application, the spliced features obtained in the step S3108 and the BERT features of the keywords obtained in the step S3208 can be input into the self-attention mechanism for feature fusion, and the query intention classifier in the self-attention mechanism can classify the fused features, so as to determine whether the commodity played in the live video is the commodity explained by the host.
[0134] In the technical solution provided in the above step S3300 of the present application, the spliced features obtained in the step S3108 and the BERT features of the keywords obtained in the step S3208 can be input into the self-attention mechanism for feature fusion, and the query intention classifier in the self-attention mechanism can classify the fused features, so as to determine whether the commodity played in the live video is the commodity explained by the host.
[0135] The present application will be further described below by a query intention (Query) extraction method.
[0136] Figure 4 is a schematic diagram of a query intention (Query) extraction method according to an embodiment of the present application. As shown in Figure 4 The method achieves extraction of query intention by processing images and audio in a live video.
[0137] In this embodiment, when processing images in a live video, first, video frames are extracted from the video, then the images in the video frames are detected and tracked to obtain entity tracking regions, and then the convolutional neural network algorithm is used to recognize the obtained entity tracking regions to obtain features such as commodity trajectory information, region of interest coordinates, and human body key points.
[0138] In this embodiment, when processing audio in a live video, first, the audio information in which the host describes the commodity is recognized by voice recognition to convert the audio information into text information, such as “baby”, “newly arrived”, “red”, “short”, “cotton-padded jacket”, etc. Then, the text information is segmented and the parts of speech are extracted to obtain the keywords of the commodity, and then the BERT encoding is used to extract the features of the keywords to obtain the BERT features of the commodity.
[0139] In this embodiment, after obtaining the product's trajectory information, region of interest coordinates, human key points, and other features, as well as the product's BERT features, these multimodal features can be input into the self-attention mechanism module for fusion to obtain the fusion result. Then, the fusion result is categorized by query intent classifier to determine whether the product played in the live video is the product being explained by the anchor.
[0140] This application significantly improves efficiency in determining whether a product played in a live video is the product being explained by the host by incorporating a self-attention mechanism. It is also relatively easy to understand intuitively. Under the effect of the self-attention mechanism, when inferring products based on their text information and trajectory features, the specific location of the entity can be forcibly determined, thereby more effectively classifying the area to which the entity belongs, and thus better classifying the products being explained by the host.
[0141] The following section will further introduce this application through application scenarios.
[0142] Figure 5 This is a schematic diagram illustrating a multimedia information processing scenario according to an embodiment of the present invention. Figure 5 As shown, the computer device plays a video and acquires multimedia information from the video, which may include image and audio information, such as live image and audio information, and then inputs it into the computing device.
[0143] In this embodiment, multimedia information in a video can be identified in a computer device to obtain information about the target object in different modalities. For example, information about the target object in the voice modal can include information about a broadcaster describing or recommending a product by voice; information about the target object in the video modal can include information about a product displayed by video; information about the target object in the image modal can include information about a product displayed by images; and information about the target object in the text modal can include information about a product described or recommended by text. This information is then output.
[0144] In this embodiment, after identifying multimedia information in a video to obtain information about the target object in different modalities, a self-attention mechanism-based fusion method is used to perform multimodal fusion of the target object's information in different modalities, thereby generating recommended content to represent the target object. For example, a feature set in the image modality can be generated based on the visual feature information of the target object, or a feature set in the text modality can be generated based on the text feature information of the target object. Then, the feature set in the image modality and the feature set in the text modality are fused to generate recommended content for the target object, which is then output to the display interface of the computing device for display.
[0145] In the prior art, the video used in the live broadcast process generally contains image information, text and audio and other multi-modal information. Due to the high complexity of the multi-modal information, the efficiency of identifying the key recommended content in the video is low. The embodiment adopts a fusion method of a self-attention mechanism to perform multi-modal fusion on the information of the target object in different modalities, and then adopts a query intention classifier to classify the fusion result, so as to improve the efficiency of classifying the goods explained by the host, thereby improving the efficiency of identifying the key recommended content in the video.
[0146] It should be noted that, for each of the foregoing method embodiments, in order to simply describe, each is described as a series of action combinations, but those skilled in the art should know that the present application is not limited to the action sequence described, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0147] From the above description of the embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and the necessary general hardware platform, and of course it can also be realized by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes a plurality of instructions for causing an end device (which can be a mobile phone, computer, server, or network device, etc.) to execute the method described in each embodiment of the present application.
[0148] Embodiment 3
[0149] According to the embodiments of the present application, a multimedia information processing apparatus for implementing the multimedia information processing method described above is also provided. It should be noted that the multimedia information processing apparatus of this embodiment can be used to execute the multimedia information processing method of the present application Figure 2 as shown.
[0150] Figure 6 is a schematic diagram of a multimedia information processing apparatus according to an embodiment of the present application. As Figure 6 shown, the multimedia information processing apparatus 60 can include a playing unit 61, an identifying unit 62 and a fusion unit 63.
[0151] The playing unit 61 is configured to play a video and acquire multimedia information in the video, wherein the multimedia information includes image information and audio information.
[0152] The identifying unit 62 is configured to identify information of the target object in different modalities from the multimedia information, wherein the target object is a product recommended in the video.
[0153] The fusing unit 63 is configured to fuse the information of the target object in different modalities to generate the recommendation content for representing the target object.
[0154] It should be noted that the playing unit 61, the identifying unit 62 and the fusing unit 63 correspond to steps S202 to S206 in Embodiment 1, and the four units have the same instances and application scenarios as the corresponding steps, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above units can run in the computer terminal 10 provided in Embodiment 1 as part of the device.
[0155] In the multimedia information processing device of the embodiment, by obtaining the multimedia information in the video, the information of the target object in different modalities in the multimedia information can be identified, so that the information of the target object in different modalities can be fused in multiple modalities, so that the recommendation content for representing the target object can be generated, thereby avoiding the problem of low efficiency in identifying the key recommendation content in the video due to the high complexity of the multi-modal information, and further solving the technical problem of low efficiency in identifying the key recommendation content in the video, and achieving the technical effect of improving the efficiency of identifying the key recommendation content in the video.
[0156] Embodiment 4
[0157] The embodiments of the present application can provide a computer terminal which can be arranged in the multimedia information processing system of the embodiments of the present application. The computer terminal can be any one of the computer terminal devices in the computer terminal group. Alternatively, in the embodiment, the computer terminal can also be replaced by a terminal device such as a mobile terminal.
[0158] Alternatively, in the embodiment, the computer terminal can be located in at least one of the network devices in the computer network.
[0159] In the embodiment, the computer terminal can execute the program codes of the following steps in the multimedia information processing method: playing a video, obtaining multimedia information in the video, wherein the multimedia information includes image information and audio information; identifying information of a target object in different modalities from the multimedia information, wherein the target object is a product recommended in the video; and fusing the information of the target object in different modalities to generate the recommendation content for representing the target object.
[0160] Alternatively, Figure 7is a structural block diagram of a mobile terminal according to an embodiment of the present application. As shown in Figure 7 The mobile terminal A can include one or more (only one is shown in the figure) processors 702, a memory 704 and a transmission device 706.
[0161] The memory can be used to store software programs and modules, such as program instructions / modules corresponding to the multimedia information processing method and device in the embodiments of the present application. The processor executes various functions and data processing by running the software programs and modules stored in the memory, i.e. implements the multimedia information processing method described above. The memory can include a high-speed random access memory, and can further include a non-volatile memory such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory can further include a memory remotely arranged with respect to the processor, which can be connected to the mobile terminal A through a network. Examples of the network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0162] The processor can call the information and application programs stored in the memory through the transmission device to perform the following steps: playing a video, obtaining multimedia information in the video, wherein the multimedia information includes image information and audio information; identifying information of a target object in different modalities from the multimedia information, wherein the target object is a recommended product in the video; and performing multi-modal fusion on the information of the target object in different modalities to generate recommendation content for representing the target object.
[0163] Optionally, the processor can further execute program codes of the following steps: image sampling on the played video to obtain image information, wherein the image information includes a video frame sequence composed of image frames; and using a detector to detect the video frame sequence to obtain visual feature information of the video on a visual track in the playing process, wherein the visual feature information includes image dimension information of the target object in different dimensions.
[0164] Optionally, the processor can further execute program codes of the following steps: detecting the video frame sequence to obtain bounding box information of at least one bounding box in the video; and identifying image dimension information of the target object played in the video based on the bounding box information of the bounding box, wherein the image dimension information includes at least one of the following: trajectory information of the target object displayed in the video frame sequence, position coordinates of the target object in each video frame, and feature information of the target object.
[0165] Optionally, the processor can further execute program codes of the following steps: performing audio sampling on the played video to obtain audio information, wherein the audio information comprises an audio frame sequence composed of audio frames; and converting the audio information into text information, wherein the text information describes the character feature information of the target object in the video.
[0166] Optionally, the processor can further execute program codes of the following steps: performing word segmentation processing on the text information to obtain at least one keyword for describing the target object; and determining the content and the part of speech of the keyword.
[0167] Optionally, the processor can further execute program codes of the following steps: generating a feature set in an image modality based on the visual feature information of the target object; generating a feature set in a text modality based on the character feature information of the target object; and generating the recommended content of the target object based on the feature set in the image modality and the feature set in the text modality.
[0168] Optionally, the processor can further execute program codes of the following steps: processing the visual feature information through a neural network model to obtain the feature set in the image modality.
[0169] Optionally, the processor can further execute program codes of the following steps: encoding the visual feature information through BERT encoding to obtain the feature set in the text modality.
[0170] Optionally, the processor can further execute program codes of the following steps: fusing the feature set in the image modality and the feature set in the text modality to generate the recommended content of the target object.
[0171] Optionally, the processor can further execute program codes of the following steps: the video is a process in which a host explains the target object; and in the fusion process, whether the target object played is the object explained by the host is determined by judging whether the feature set in the image modality and the feature set in the text modality have a corresponding relationship.
[0172] As another optional example, the processor can call the information and the application program stored in the memory through the transmission device to execute the following steps: inputting multimedia information in the played video in an input interface of an operation interface, wherein the multimedia information comprises image information and audio information; sensing a recommended content generation instruction in the operation interface, and identifying information of a target object in different modalities from the multimedia information, wherein the target object is a product recommended in the video; and displaying the recommended content for representing the target object on the operation interface, wherein the recommended content of the target object is generated by performing multi-modal fusion on the information of the target object in different modalities.
[0173] As another optional example, the processor can call information and application programs stored in the memory through the transmission device to perform the following steps: displaying the played video and multimedia information in the video in the operation interface, wherein the multimedia information includes image information and audio information; sensing a recommendation content generation instruction in the operation interface; in response to the content generation instruction, displaying information of a target object in different modalities on the operation interface, wherein the target object is a product recommended in the video, and the information of the target object in different modalities is identified from the multimedia information; and displaying recommendation content for representing the target object on the operation interface, wherein the recommendation content of the target object is generated by performing multi-modal fusion on the information of the target object in different modalities.
[0174] As another optional example, the processor can call information and application programs stored in the memory through the transmission device to perform the following steps: uploading multimedia information in the played video by the front-end client, wherein the multimedia information includes image information and audio information; transmitting the multimedia information to the background server by the front-end client; receiving information of a target object in different modalities returned by the background server from the multimedia information by the front-end client, wherein the target object is a product recommended in the video; and performing multi-modal fusion on the information of the target object in different modalities by the front-end client to generate recommendation content for representing the target object.
[0175] As another optional example, the processor can call information and application programs stored in the memory through the transmission device to perform the following steps: receiving a recommendation content generation request; obtaining multimedia information of a video in the recommendation content generation request, wherein the multimedia information includes image information and audio information; identifying information of a target object in different modalities from the multimedia information, wherein the target object is a product recommended in the video; performing multi-modal fusion on the information of the target object in different modalities to generate recommendation content for representing the target object; and outputting the recommendation content of the target object.
[0176] The embodiment of the present application provides a scheme of a multimedia information processing method. Video is played, and multimedia information in the video is acquired, wherein the multimedia information comprises image information and audio information; information of a target object in different modalities is identified from the multimedia information, wherein the target object is a recommended product in the video; and information of the target object in different modalities is fused in a multimodal manner to generate recommended content for representing the target object, that is, information of the target object in different modalities in the multimedia information can be identified by acquiring the multimedia information in the video, so that the information of the target object in different modalities can be fused in a multimodal manner, and thus the recommended content for representing the target object can be generated, thereby avoiding the problem of low efficiency in identifying key recommended content in the video due to high complexity of multimodal information, and further solving the technical problem of low efficiency in identifying key recommended content in the video, and achieving the technical effect of improving the efficiency in identifying key recommended content in the video.
[0177] Those skilled in the art can understand that, Figure 7 The structure shown is only schematic, and the mobile terminal A can also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, a mobile Internet device (MID), a PAD, or the like. Figure 7 It does not limit the structure of the electronic device. For example, the mobile terminal A can further include more or fewer components (such as a network interface, a display device, etc.) than those shown in the figure, or have a different configuration from that shown in the figure. Figure 7 It does not limit the structure of the electronic device. For example, the mobile terminal A can further include more or fewer components (such as a network interface, a display device, etc.) than those shown in the figure, or have a different configuration from that shown in the figure. Figure 7 It does not limit the structure of the electronic device. For example, the mobile terminal A can further include more or fewer components (such as a network interface, a display device, etc.) than those shown in the figure, or have a different configuration from that shown in the figure.
[0178] Those skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by programs instructing the related hardware of the terminal device, and the programs can be stored in a computer readable storage medium, which can include a flash disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0179] Embodiment 5
[0180] The embodiment of the present application further provides a computer readable storage medium. Optionally, in the embodiment, the storage medium can be used to save the program code executed by the multimedia information processing method provided in the embodiment 1.
[0181] Optionally, in the embodiment, the storage medium can be located in any one of the computer terminals in a computer terminal group in a computer network, or in any one of the mobile terminals in a mobile terminal group.
[0182] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: playing a video, acquiring multimedia information in the video, wherein the multimedia information includes: image information and audio information; identifying information of a target object in different modalities from the multimedia information, wherein the target object is a product recommended in the video; and performing multimodal fusion of the information of the target object in different modalities to generate recommended content for representing the target object.
[0183] Optionally, the storage medium is also configured to store program code for performing the following steps: sampling the played video to obtain image information, wherein the image information includes a video frame sequence composed of image frames; detecting the video frame sequence using a detector to obtain visual feature information on the visual trajectory of the video during playback, wherein the visual feature information includes image dimension information of the target object in different dimensions.
[0184] Optionally, the storage medium is further configured to store program code for performing the following steps: detecting a video frame sequence and obtaining bounding box information of at least one bounding box in the video; based on the bounding box information of the bounding boxes, identifying image dimension information of a target object playing in the video, wherein the image dimension information includes at least one of the following: trajectory information of the target object displayed in the video frame sequence, position coordinates of the target object in each video frame, and feature information of the target object.
[0185] Optionally, the storage medium is also configured to store program code for performing the following steps: sampling the audio of the played video to obtain audio information, wherein the audio information includes an audio frame sequence consisting of audio frames; and converting the audio information into text information, wherein the text information describes the textual feature information of the target object played in the video.
[0186] Optionally, the storage medium is also configured to store program code for performing the following steps: segmenting text information to obtain at least one keyword for describing the target object; and determining the content and part of speech of the keyword.
[0187] Optionally, the storage medium is also configured to store program code for performing the following steps: generating a feature set in the image modality based on the visual feature information of the target object; generating a feature set in the text modality based on the text feature information of the target object; and generating recommended content for the target object based on the feature sets in the image modality and the feature sets in the text modality.
[0188] Optionally, the storage medium is also configured to store program code for performing the following steps: processing visual feature information through a neural network model to obtain a feature set in the image modality.
[0189] Optionally, the storage medium is further configured to store program code for encoding the visual feature information by BERT encoding to obtain the feature set in the text mode.
[0190] Optionally, the storage medium is further configured to store program code for fusing the feature set in the image mode and the feature set in the text mode to generate the recommended content of the target object.
[0191] Optionally, the storage medium is further configured to store program code for performing the following steps: if the video is a process of the anchor explaining the target object, then in the fusion process, it is determined whether the target object played is the object explained by the anchor by judging whether the feature set in the image mode and the feature set in the text mode have a corresponding relationship.
[0192] As an optional example, the computer readable storage medium is further configured to store program code for performing the following steps: entering multimedia information in the played video in the input interface of the operation interface, wherein the multimedia information includes image information and audio information; sensing the recommended content generation instruction in the operation interface, and identifying information of the target object in different modes from the multimedia information, wherein the target object is a product recommended in the video; and displaying the recommended content for representing the target object on the operation interface, wherein the recommended content of the target object is generated by multi-modal fusion of the information of the target object in different modes.
[0193] As an optional example, the computer readable storage medium is further configured to store program code for performing the following steps: displaying the played video in the operation interface, and displaying multimedia information in the video, wherein the multimedia information includes image information and audio information; sensing the recommended content generation instruction in the operation interface; in response to the content generation instruction, displaying information of the target object in different modes on the operation interface, wherein the target object is a product recommended in the video, and the information of the target object in different modes is identified from the multimedia information; and displaying the recommended content for representing the target object on the operation interface, wherein the recommended content of the target object is generated by multi-modal fusion of the information of the target object in different modes.
[0194] As an optional example, the computer readable storage medium is further configured to store program code for performing the following steps: the front-end client uploads multimedia information in a played video, wherein the multimedia information comprises image information and audio information; the front-end client transmits the multimedia information to the back-end server; the front-end client receives information about a target object in different modalities returned by the back-end server, wherein the target object is a recommended product in the video; and the front-end client performs multi-modal fusion on the information about the target object in different modalities to generate recommendation content for representing the target object.
[0195] As an optional example, the computer readable storage medium is further configured to store program code for performing the following steps: receiving a recommendation content generation request; obtaining multimedia information of a video in the recommendation content generation request, wherein the multimedia information comprises image information and audio information; identifying information about a target object in different modalities from the multimedia information, wherein the target object is a recommended product in the video; performing multi-modal fusion on the information about the target object in different modalities to generate recommendation content for representing the target object; and outputting the recommendation content of the target object.
[0196] The above-mentioned serial numbers of the embodiments of the application are only for description, and do not represent the advantages or disadvantages of the embodiments.
[0197] In the above-mentioned embodiments of the application, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.
[0198] In several embodiments provided in the present application, it should be understood that the disclosed technical contents can be implemented by other manners. Among them, the above-mentioned device embodiments are only schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed units can be indirect coupling or communication connection through some interfaces, units or modules, which can be electrical or other forms.
[0199] The units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. According to actual needs, part or all of the units can be selected to achieve the purpose of the embodiment scheme.
[0200] In addition, each function unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function unit.
[0201] When the integrated unit is realized in the form of a software function unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application, essentially or the part that contributes to the prior art, or all or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.
[0202] The above is only the preferred embodiment of the present application, and it should be pointed out that for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements should be considered as the protection scope of the present application.
Claims
1. A method for processing multimedia information, characterized in that, include: Play a video and acquire multimedia information from the video, wherein the multimedia information includes: image information and audio information; Information about a target object in different modalities is identified from the multimedia information, wherein the target object is a product recommended in the video; The information of the target object in different modalities is fused in a multimodal manner to generate recommended content that represents the target object; The process of multimodal fusion of information about the target object in different modalities to generate recommended content representing the target object includes: generating a feature set in the image modality using a convolutional neural network based on the visual feature information of the target object; generating a feature set in the text modality using BERT encoding based on the text feature information of the target object; inputting the feature sets in the image modality and the feature sets in the text modality into a self-attention mechanism model for fusion to obtain a fusion result; classifying the fusion result using a query intent classifier to obtain a classification result, wherein the classification result is used to indicate whether the target object in the image information is the target object in the audio information; and generating the recommended content based on the classification result.
2. The method according to claim 1, characterized in that, Identifying information about the target object in different modalities from the multimedia information includes: The video being played is sampled to obtain image information, wherein the image information includes a video frame sequence composed of image frames; The video frame sequence is detected using a detector to obtain visual feature information on the visual trajectory of the video during playback, wherein the visual feature information includes image dimensional information of the target object in different dimensions.
3. The method according to claim 2, characterized in that, The video frame sequence is detected using a detector to obtain visual feature information on the visual trajectory of the video during playback, including: Detect the video frame sequence and obtain the bounding box information of at least one bounding box in the video; Based on the bounding box information of the bounding box, the image dimension information of the target object playing in the video is identified, wherein the image dimension information includes at least one of the following: the trajectory information of the target object displayed in the video frame sequence, the position coordinates of the target object in each video frame, and the feature information of the target object.
4. The method according to claim 2 or 3, characterized in that, Identifying information about the target object in different modalities from the multimedia information includes: The video being played is sampled for audio to obtain audio information, wherein the audio information includes an audio frame sequence consisting of audio frames; The audio information is converted into text information, wherein the text information describes the textual feature information of the target object played in the video.
5. The method according to claim 4, characterized in that, After converting the audio information into text information, the method further includes: The text information is segmented to obtain at least one keyword for describing the target object; Determine the content and part of speech of the keywords.
6. The method according to claim 1, characterized in that, The visual feature information is processed by a neural network model to obtain the feature set of the image modality.
7. The method according to claim 1, characterized in that, The visual feature information is encoded using BERT encoding to obtain the feature set for the text modality.
8. The method according to claim 1, characterized in that, The feature sets of the image modality and the feature sets of the text modality are fused to generate recommended content for the target object.
9. The method according to claim 8, characterized in that, If the video is a process of the anchor explaining the target object, then during the fusion process, the target object being played is determined to be the object being explained by the anchor by judging whether there is a corresponding relationship between the feature set in the image modality and the feature set in the text modality.
10. A method for processing multimedia information, characterized in that, include: The multimedia information from the video being played is entered in the input interface of the operation interface, wherein the multimedia information includes: image information and audio information; Within the operation interface, a recommended content generation instruction is sensed, and the information of the target object in different modalities is identified from the multimedia information, wherein the target object is the product recommended in the video; The operation interface displays recommended content representing the target object, wherein the recommended content is generated based on classification results. The classification results indicate whether the target object in the image information is the same as the target object in the audio information. The classification results are obtained by performing intent classification on the fusion results through a query intent classifier. The fusion results are obtained by inputting the feature sets in the image modality and the feature sets in the text modality into a self-attention mechanism model for fusion. The feature sets in the image modality are obtained using a convolutional neural network based on the visual feature information of the target object, and the feature sets in the text modality are obtained using BERT encoding based on the text feature information of the target object.
11. A method for processing multimedia information, characterized in that, include: The operating interface displays the playing video and the multimedia information in the video, wherein the multimedia information includes: image information and audio information; The recommended content generation command is sensed within the user interface. In response to the content generation instruction, information about the target object in different modalities is displayed on the operation interface. The target object is a product recommended in the video. The information about the target object in different modalities is identified from the multimedia information. The recommended content is generated based on classification results. The classification results indicate whether the target object in the image information is the same as the target object in the audio information. The classification results are obtained by performing intent classification on the fusion results using a query intent classifier. The fusion results are obtained by inputting the feature sets from the image modality and the feature sets from the text modality into a self-attention mechanism model for fusion. The feature sets from the image modality are obtained using a convolutional neural network based on the visual feature information of the target object, and the feature sets from the text modality are obtained using BERT encoding based on the text feature information of the target object. The recommended content for representing the target object is displayed on the operation interface. The recommended content for the target object is generated by multimodal fusion of information of the target object in different modalities.
12. A method for processing multimedia information, characterized in that, include: The multimedia information in the video uploaded and played by the front-end client includes: image information and audio information; The front-end client transmits the multimedia information to the back-end server; The front-end client receives information from the back-end server about the target object identified from the multimedia information in different modalities. The target object is a product recommended in the video. The recommended content is generated based on classification results. The classification results indicate whether the target object in the image information is the same as the target object in the audio information. The classification results are obtained by performing intent classification on the fusion results using a query intent classifier. The fusion results are obtained by inputting the feature sets from the image modality and the feature sets from the text modality into a self-attention mechanism model for fusion. The feature sets from the image modality are obtained using a convolutional neural network based on the visual feature information of the target object, and the feature sets from the text modality are obtained using BERT encoding based on the text feature information of the target object. The front-end client performs multimodal fusion of information about the target object in different modalities to generate recommended content that represents the target object.
13. A method for processing multimedia information, characterized in that, include: Receive a request to generate recommended content; Obtain multimedia information of the video in the recommended content generation request, wherein the multimedia information includes: image information and audio information; Information about a target object in different modalities is identified from the multimedia information, wherein the target object is a product recommended in the video; The information of the target object in different modalities is fused in a multimodal manner to generate recommended content that represents the target object; Output the recommended content for the target object; The process of multimodal fusion of information about the target object in different modalities to generate recommended content representing the target object includes: generating a feature set in the image modality using a convolutional neural network based on the visual feature information of the target object; generating a feature set in the text modality using BERT encoding based on the text feature information of the target object; inputting the feature sets in the image modality and the feature sets in the text modality into a self-attention mechanism model for fusion to obtain a fusion result; classifying the fusion result using a query intent classifier to obtain a classification result, wherein the classification result is used to indicate whether the target object in the image information is the target object in the audio information; and generating the recommended content based on the classification result.
14. A multimedia information processing device, characterized in that, include: A playback unit is used to play a video and acquire multimedia information from the video, wherein the multimedia information includes image information and audio information; The identification unit is used to identify information about a target object in different modalities from the multimedia information, wherein the target object is a product recommended in the video; The fusion unit is used to perform multimodal fusion of information of the target object under different modalities to generate recommended content to represent the target object; The fusion unit is configured to obtain the recommended content through the following steps: using a convolutional neural network, generating a feature set in the image modality based on the visual feature information of the target object; using BERT encoding, generating a feature set in the text modality based on the text feature information of the target object; inputting the feature set in the image modality and the feature set in the text modality into a self-attention mechanism model for fusion to obtain a fusion result; using a query intent classifier to classify the fusion result to obtain a classification result, wherein the classification result is used to indicate whether the target object in the image information is the target object in the audio information; and generating the recommended content based on the classification result.
15. A live streaming all-in-one machine, characterized in that, include: A player is used to play videos and acquire multimedia information from the videos, wherein the multimedia information includes image information and audio information; The recognizer, connected to the player, is used to identify information about a target object in different modalities from the multimedia information, wherein the target object is a product recommended in the video; A generator, connected to the recognizer, is used to perform multimodal fusion of information of the target object in different modalities to generate recommended content to represent the target object; The generator is configured to obtain the recommended content through the following steps: using a convolutional neural network, generating a feature set in the image modality based on the visual feature information of the target object; using BERT encoding, generating a feature set in the text modality based on the text feature information of the target object; inputting the feature set in the image modality and the feature set in the text modality into a self-attention mechanism model for fusion to obtain a fusion result; using a query intent classifier to classify the fusion result to obtain a classification result, wherein the classification result is used to indicate whether the target object in the image information is the target object in the audio information; and generating the recommended content based on the classification result.
16. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device on which the computer-readable storage medium is located to perform the method according to any one of claims 1 to 13.
17. A processor, characterized in that, The processor is used to run a program, wherein the program, when running, performs the method according to any one of claims 1 to 13.
18. A multimedia information processing system, characterized in that, include: processor; A memory, connected to the processor, is configured to provide the processor with instructions to perform the following processing steps: playing a video; acquiring multimedia information from the video, wherein the multimedia information includes image information and audio information; identifying information of a target object in different modalities from the multimedia information, wherein the target object is a product recommended in the video; and performing multimodal fusion of the information of the target object in different modalities to generate recommended content characterizing the target object. The memory is further configured to provide the processor with instructions to perform the following processing steps: using a convolutional neural network to generate a feature set in the image modality based on the visual feature information of the target object; using BERT encoding to generate a feature set in the text modality based on the text feature information of the target object; inputting the feature set in the image modality and the feature set in the text modality into a self-attention mechanism model for fusion to obtain a fusion result; using a query intent classifier to classify the fusion result to obtain a classification result, wherein the classification result is used to indicate whether the target object in the image information is the target object in the audio information; and generating the recommended content based on the classification result.
Citation Information
Patent Citations
Commodity recommendation method and device based on video
CN106202317A
Live broadcast system, method, apparatus, and electronic device for determining live video theme
CN109104639A
Video classification method and device and electronic equipment
CN110399934A
Topic classification method, device and equipment based on multiple modes and storage medium
CN111259215A