Regional attribute identification method and related device
Through cross-modal feature fusion and global feature extraction, the accuracy of multimedia resource regional attribute recognition in the prior art is solved, and the precise recognition of the regionality of multimedia resource and the accurate positioning of regional names are achieved.
Patent Information
- Application Number
- CN202510510612.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-07-22
AI Technical Summary
In the prior art, the regional attribute recognition method based on multimedia resources relies on the extraction of feature of partial text content, resulting in the inability to fully and accurately identify the regionality and regional names of the multimedia resources. Especially when the text content lacks regional information or is unclear, the identification accuracy is limited.
The cross-modal feature fusion method is used to extract the initial features of the text and non-text modal data of multimedia resources, and the target region name associated with the regional event is identified through cross-modal feature fusion and global feature extraction of non-text modal modes.
The accuracy of multimedia resource regional recognition is improved, and the real region name of the multimedia resource can be determined more accurately, and the recognition error caused by insufficient text content in the prior art is overcome.
Smart Images

Figure CN120354362A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and provides a geographical attribute recognition method and related device. Background Art
[0002] With the continuous development of multimedia technology, more and more application scenarios require the recognition of the geographical attributes of multimedia resources; among them, the geographical attribute is used to indicate whether the multimedia resource has geographical characteristics, and when the multimedia resource has geographical characteristics, the geographical attribute is also used to indicate the geographical name associated with the geographical event.
[0003] For example, in a news application, it is necessary to determine whether a news has geographical characteristics, and when the news has geographical characteristics, determine the geographical name associated with the geographical event, and then display the news with a specific geographical name in the geographical recommendation channel;
[0004] For another example, in hot event analysis, it is necessary to determine whether the articles related to a hot event have geographical characteristics, and when they have geographical characteristics, determine the geographical name associated with the geographical hot event, so as to know the association between the hot event and the geographical name.
[0005] In the related art, usually based on a part of the text content selected from the multimedia resource, feature extraction is performed on the selected text content for geographical recognition, and geographical name extraction is performed on the selected text content. However, the selected part of the text content often contains limited descriptive information and cannot comprehensively and accurately describe the region. When the selected text content does not contain region-related information, or there is unclear region information, it will seriously affect the accuracy of geographical attribute recognition.
[0006] For example, assume that the news title 1 is "The production process of A food". In news title 1, there is no region-related information, but the main content of this news is to introduce the A food produced in place B. Then, just through news title 1, this news cannot be recognized as having geographical characteristics.
[0007] For another example, assume that the news title 2 is "Joint activities of universities B, C, and D". Among them, B is a geographical name, and C and D are not geographical names. In news title 2, although it contains the region-related information "B", the main content of this news is to introduce the joint activities between universities, and the place where the joint activities occur is not in place B. Then, just through news title 2, although this news is recognized as having geographical characteristics, the recognized geographical name is not the geographical name associated with the geographical event.
[0008] In view of this, it is necessary to provide a new geographical attribute recognition method to overcome the above defects. Summary of the Invention
[0009] An embodiment of the present application provides a geographical attribute recognition and related device for accurately recognizing the geographical attributes of multimedia resources.
[0010] On the one hand, an embodiment of the present application provides a geographical attribute recognition method, including:
[0011] Feature extraction is respectively performed on at least two types of modal data extracted from the multimedia resource to obtain corresponding initial features; the at least two types of modal data include: text modal data and at least one non-text modal data; the text modal data includes at least one of the following: at least one initial text data in the multimedia resource, and transformed text data obtained by modal transformation of at least one non-text modal initial non-text data in the multimedia resource;
[0012] Cross-modal feature fusion is performed on the obtained at least two initial features to obtain a fused feature;
[0013] For the initial feature of each non-text modality, the following operations are respectively performed: global feature extraction is respectively performed on multiple sub-features in an initial feature to obtain corresponding non-text global features; each sub-feature represents: a sub-data item in the corresponding non-text modal data;
[0014] Based on the fused feature, in combination with the obtained at least one non-text global feature, a geographical recognition result of the multimedia resource is obtained;
[0015] When the geographical recognition result indicates that the multimedia resource has geographical characteristics, based on the text modal data, the target geographical name associated with the geographical event corresponding to the multimedia resource is recognized.
[0016] On the one hand, an embodiment of the present application provides a geographical attribute recognition device, including:
[0017] A feature extraction module, configured to respectively perform feature extraction on at least two types of modal data extracted from the multimedia resource to obtain corresponding initial features; the at least two types of modal data include: text modal data and at least one non-text modal data; the text modal data includes at least one of the following: at least one initial text data in the multimedia resource, and transformed text data obtained by modal transformation of at least one non-text modal initial non-text data in the multimedia resource;
[0018] A feature fusion module, configured to perform cross-modal feature fusion on the obtained at least two initial features to obtain a fused feature;
[0019] A non-text processing module, which is used to respectively execute for each initial feature of each non-text modality: globally extract features from multiple sub-features in an initial feature to obtain corresponding non-text global features; each sub-feature represents: a piece of sub-data in the corresponding non-text modality data.
[0020] A regional recognition module, which is used to obtain a regional recognition result of the multimedia resource based on the fusion feature and in combination with at least one obtained non-text global feature.
[0021] A regional name recognition module, which is used to, when the regional recognition result indicates that the multimedia resource has regionality, identify a target regional name associated with the regional event corresponding to the multimedia resource based on the text modality data.
[0022] In some optional embodiments, the feature fusion module is specifically used for:
[0023] Align the dimensions of the at least two initial features;
[0024] For each initial feature after dimension alignment, respectively execute: in an initial feature, add a modality feature element and a position feature element to obtain a corresponding target feature; the modality feature element is used to describe the modality of the modality data corresponding to the one initial feature; the position feature element is used to describe the positions of the respective sub-data included in the modality data corresponding to the one initial feature in the modality data.
[0025] Perform cross-modal feature fusion based on the at least two obtained target features to obtain the fusion feature.
[0026] In some optional embodiments, the feature fusion module is specifically used for:
[0027] Perform splicing based on the at least two target features to obtain a first spliced feature;
[0028] For each feature element in the first spliced feature, respectively execute: respectively determine the attention weights between a feature element and each feature element in the first spliced feature, and each attention weight represents the dependency relationship between two feature elements; based on the respective feature elements and the respective attention weights corresponding to the respective feature elements, obtain a fusion element corresponding to the one feature element.
[0029] Obtain the fusion feature based on the obtained multiple fusion elements.
[0030] In some optional embodiments, the regional recognition module is specifically used for:
[0031] Align the dimensions of the fusion feature and the at least one non-text global feature, and then splice them to obtain a second spliced feature;
[0032] Perform a classification mapping on the second spliced feature to obtain target classification information;
[0033] Use the region recognition result corresponding to the target classification information as the region recognition result of the multimedia resource.
[0034] In some optional embodiments, the region name recognition module is specifically configured to:
[0035] Perform an associated region analysis of region events based on the text modality data to obtain an original region name, where the original region name includes at least one original region word segment;
[0036] Based on the respective original region levels corresponding to the at least one original region word segment, and in combination with a preset region hierarchy relationship, perform a recognition of the missing region level for the original region name to obtain a recognition result;
[0037] When the recognition result indicates that the original region name does not lack a region level to be filled, use the original region name as the target region name;
[0038] When the recognition result indicates that the original region name lacks at least one region level to be filled, from the region word segment relationship library, for each region level to be filled, obtain the region word segment to be filled associated with the at least one original region word segment, and respectively fill each obtained region word segment to be filled into the original region name to obtain the target region name.
[0039] In some optional embodiments, the region name recognition module is specifically configured to:
[0040] When the text modality data includes one item of data, perform an associated region analysis of region events on the one item of data to obtain the original region name; the one item of data is: one item of initial text data, or, transformed text data corresponding to a non-text modality;
[0041] When the text modality data includes multiple items of data, based on the respective priorities of the multiple items of data, sequentially perform an associated region analysis of region events on the multiple items of data to obtain the original region name; the multiple items of data are any one of the following: multiple items of initial text data; transformed text data corresponding to multiple non-text modalities; at least one item of initial text data and transformed text data corresponding to at least one non-text modality.
[0042] In some optional embodiments, the region name recognition module is specifically configured to:
[0043] Read the multiple pieces of data in order of priority. For each piece of data read, perform the following operations:
[0044] Perform an associated region analysis of regional events on the piece of data to obtain at least one regional word segment;
[0045] When the piece of data is the first piece of data read, use the at least one regional word segment as the regional name corresponding to the piece of data;
[0046] When the piece of data is not the first piece of data read, supplement the regional name corresponding to the previous piece of data read based on the at least one regional word segment to obtain the regional name corresponding to the piece of data;
[0047] When the regional name corresponding to the piece of data contains a regional word segment at the target level in the regional hierarchy, or when the piece of data is the last piece of data read, obtain the original regional name based on the regional name corresponding to the piece of data.
[0048] On the one hand, an embodiment of the present application provides a computer device, including a processor and a memory. The memory stores a computer program. When the computer program is executed by the processor, the processor is caused to execute the steps of the above method.
[0049] On the one hand, an embodiment of the present application provides a computer-readable storage medium, which includes a computer program. When the computer program runs on a computer device, the computer program is used to cause the computer device to execute the steps of the method in any of the above aspects.
[0050] On the one hand, an embodiment of the present application provides a computer program product. The program product includes a computer program. The computer program is stored in a computer-readable storage medium. A processor of a computer device reads and executes the computer program from the computer-readable storage medium, causing the computer device to execute the steps of the method in any of the above aspects.
[0051] In the embodiments of the present application, for at least two types of modal data corresponding to multimedia resources, first, feature extraction is respectively performed on the at least two types of modal data to obtain corresponding initial features, and then cross-modal feature fusion is performed on the obtained at least two initial features to obtain fusion features. By performing cross-modal feature fusion on the initial features of multiple modalities, the interaction information between multiple modalities of data is captured. This interaction information reflects the mutual dependence between different modalities of data and helps subsequent regional recognition.
[0052] Further, for the initial features of each non-text modality, global feature extraction is performed on multiple sub-features included in an initial feature to obtain corresponding non-text global features. Since the cross-modal feature fusion process will focus more on the text modality, by performing global feature extraction on the initial features of the non-text modality, the expression of the initial features of the non-text modality is strengthened. In this way, based on the fusion features and the non-text global features, the accuracy of the region recognition result can be effectively improved.
[0053] In addition, when the region recognition result indicates that the multimedia resource has regionality, based on the text modality data, the target region name associated with the regional event corresponding to the multimedia resource is recognized. Since the target region name is associated with the regional event, compared with directly extracting the region name from the selected part of the text content, the region name truly relevant to the multimedia resource can be determined more accurately.
[0054] Other features and advantages of the present application will be described in the following specification, and part of them will become obvious from the specification, or be understood by implementing the present application. The objectives and other advantages of the present application can be achieved and obtained through the structures specifically pointed out in the written specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0056] Figure 1 is a schematic diagram of an application scenario provided in an embodiment of the present application;
[0057] Figure 2 is a schematic flowchart of a method for identifying regional attributes provided in an embodiment of the present application;
[0058] Figure 3 is a schematic diagram of the feature extraction logic of visual modality data provided in an embodiment of the present application Figure 1 ;
[0059] Figure 4 is a schematic diagram of the feature extraction logic of visual modality data provided in an embodiment of the present application Figure 2 ;
[0060] Figure 5 is a schematic diagram of the feature extraction logic of visual modality data provided in an embodiment of the present application Figure 3 ;
[0061] Figure 6 is a schematic diagram of the target features of the text modality provided in an embodiment of the present application;
[0062] Figure 7 Schematic diagram of target features of the visual modality provided in the embodiments of the present application;
[0063] Figure 8 Schematic diagram of the logic of cross-modal feature fusion provided in the embodiments of the present application;
[0064] Figure 9 Schematic diagram of the logic of regional identification provided in the embodiments of the present application;
[0065] Figure 10 Schematic diagram of the logic of the process of regional attribute identification provided in the embodiments of the present application;
[0066] Figure 11 Schematic diagram of the structure of a regional attribute identification device provided in the embodiments of the present application;
[0067] Figure 12 Schematic diagram of the structure of a computer device provided in the embodiments of the present application. Detailed implementation manners
[0068] In order to make the objectives, technical solutions and beneficial effects of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0069] Some terms in the embodiments of the present application are explained below to facilitate the understanding of those skilled in the art.
[0070] (1) The multimedia resources are various forms of articles such as pictures and texts, videos, short essays, etc.
[0071] (2) An event refers to a point in time and space specified by time and space. The description of an event in multimedia resources usually covers key element information such as time, place, person, action, etc.
[0072] (3) A regional event refers to an event associated with a specific region. Generally speaking, a regional event can have a strong connection with a specific region, have a certain impact on a specific region, or reflect the local customs and culture and image of a specific region. Such as events like people's livelihood policies, local life, social sudden disaster events, historical culture, or city business cards.
[0073] (4) The Transformer neural network is an architecture based on the attention mechanism, and through the attention mechanism, long-range dependencies in the sequence can be captured.
[0074] It is understandable that when the embodiments in the present application are applied to specific products or technologies, object permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions.
[0075] The following briefly introduces the design concept of the embodiments of the present application.
[0076] With the continuous development of multimedia technology, more and more application scenarios require the recognition of the geographical attributes of multimedia resources; among them, the geographical attribute is used to indicate whether the multimedia resource has geographical characteristics, and when the multimedia resource has geographical characteristics, the geographical attribute is also used to indicate the geographical name associated with the geographical event.
[0077] For example, in a news application, it is necessary to determine whether the news has geographical characteristics, and when the news has geographical characteristics, determine the geographical name associated with the geographical event, and then display the news with a specific geographical name in the geographical recommendation channel.
[0078] For another example, in hotspot event analysis, it is necessary to determine whether the articles related to the hotspot event have geographical characteristics, and when they have geographical characteristics, determine the geographical name associated with the geographical hotspot event, so as to know the association between the hotspot event and the geographical name.
[0079] For yet another example, in article statistics, it is necessary to determine whether the included articles have geographical characteristics, and when they have geographical characteristics, determine the geographical name associated with the included articles, so as to count the number of included articles corresponding to each geographical name.
[0080] In the related art, usually, based on a part of the text content selected from the multimedia resource, feature extraction is performed on the selected text content for geographical recognition, and geographical name extraction is performed on the selected text content (such as using a word segmentation tool to segment the selected text content and extracting place names from the segmentation results). However, the selected part of the text content often contains limited descriptive information and cannot comprehensively and accurately describe the region. When the selected text content does not contain geographical information or contains unclear geographical information, it will seriously affect the accuracy of geographical attribute recognition.
[0081] For example, the news title 1 is "The production process of A food". In news title 1, no geographical information is included, but the main content of the news is to introduce the A food produced in place B. Then, it is impossible to recognize this news as having geographical characteristics only through news title 1.
[0082] For another example, assume that news title 2 is "Joint Activity of Universities B, C, and D". Among them, B is a geographical name, while C and D are not geographical names. In news title 2, although it contains geographical information "B", the main content of this news is to introduce the joint activity among universities, and the place where the joint activity occurs is not in area B. Therefore, although this news is identified as having geographical characteristics based on news title 2 alone, the identified geographical name is not the area associated with the geographical event.
[0083] In view of this, an embodiment of the present application provides a geographical attribute recognition method. In this method, for at least two types of modal data corresponding to a multimedia resource, first, feature extraction is respectively performed on the at least two types of modal data to obtain corresponding initial features, and then cross-modal feature fusion is performed on the obtained at least two initial features to obtain a fused feature. By performing cross-modal feature fusion on the initial features of multiple modalities, the interaction information between multiple modalities of data is captured. These interaction information reflect the mutual dependence between different modalities of data, which is helpful for subsequent geographical recognition.
[0084] Furthermore, for the initial feature of each non-text modality, global feature extraction is respectively performed on multiple sub-features included in an initial feature to obtain corresponding non-text global features. Since the cross-modal feature fusion process will focus more on the text modality, by performing global feature extraction on the initial features of non-text modalities, the expression of the initial features of non-text modalities is strengthened. In this way, based on the fused feature and the non-text global features, the accuracy of the geographical recognition result can be effectively improved.
[0085] In addition, when the geographical recognition result indicates that the multimedia resource has geographical characteristics, based on the text modality data, the target geographical name associated with the geographical event corresponding to the multimedia resource is recognized. Since the target geographical name is associated with the geographical event, compared with directly extracting the geographical name from the selected part of the text content, the geographical name truly related to the multimedia resource can be determined more accurately.
[0086] Next, a simple introduction is made to the application scenarios applicable to the technical solutions of the embodiments of the present application. It should be noted that the following introduced application scenarios are only used to illustrate the embodiments of the present application rather than to limit them. In the specific implementation process, the technical solutions provided by the embodiments of the present application can be flexibly applied according to actual needs.
[0087] The solution provided by the embodiments of the present application can be applied to various scenarios that require geographical attribute recognition, such as the display scenario of the geographical recommendation channel in news applications, etc. Refer to Figure 1 As shown, it is a schematic diagram of an application scenario provided by an embodiment of the present application. In this scenario, it may include a terminal device 301 and a server 302.
[0088] In one implementation, the terminal device 301 can be, for example, a mobile phone, a tablet computer, a laptop computer, a desktop computer, a smart TV, a smart vehicle-mounted device, and a smart wearable device, etc. The terminal device 301 can install a service application, and the service application can be a software client, or a web page, a small program, etc. The server 302 is the background server corresponding to the software, web page, small program, etc., and the specific type of the client is not limited.
[0089] The server 302 can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms, but is not limited thereto.
[0090] In one implementation, the server 302 can include one or more processors, a memory, and an I / O interface for interacting with the terminal, etc. In addition, the server 302 can also be configured with a database, and the database can be used to store multi-modal data associated with the target object and trained model parameters, etc. Among them, the memory of the server 302 can also store the program instructions of the geographical attribute recognition method provided by the embodiments of the present application. When these program instructions are executed by the processor, they can be used to implement the steps of the geographical attribute recognition method provided by the embodiments of the present application to realize the geographical attribute recognition process.
[0091] In one implementation, the geographical attribute recognition method provided by the embodiments of the present application can be implemented through a geographical attribute recognition model, and the geographical attribute recognition model can include multiple network models. Specifically, it involves the training process of the geographical attribute recognition model and the application process of the geographical attribute recognition model. In the training process, the geographical attribute recognition model is iteratively trained using sample data to obtain the trained geographical attribute recognition model. In the application process, at least two types of modal data extracted from the multimedia resource are input into the trained geographical attribute recognition model to obtain the prediction result of the object category to which the target object belongs.
[0092] In one implementation, the training process of the geographical attribute recognition model can be executed by the server 302 to utilize the computing resources of the server 302 to quickly implement the training of the geographical attribute recognition model, and the application process of the geographical attribute recognition model can also be participated in and executed by the terminal device 301. For example, the terminal device 301 uses the feature extraction part in the trained geographical attribute recognition model to extract features from at least two types of modal data, obtain corresponding initial features, and transmit the obtained at least two initial features to the server 302. Furthermore, the server 302 can perform geographical attribute recognition based on the obtained at least two initial features to obtain a geographical recognition result and a target geographical name.
[0093] The terminal device 301 and the server 302 can be directly or indirectly communicatively connected through one or more networks. The network can be a wired network or a wireless network. For example, the wireless network can be a mobile cellular network or a Wireless-Fidelity (WIFI) network. Of course, it can also be other possible networks, and the embodiments of the present invention do not limit this.
[0094] It should be noted that in the embodiments of the present application, the number of terminal devices 301 can be one or multiple. Similarly, the number of servers 302 can also be one or multiple. That is to say, the number of terminal devices 301 or servers 302 is not limited.
[0095] In a possible application scenario, the relevant data (such as multimedia resources, at least two types of modal data, initial features, fusion features, and non-text global features, etc.) and model parameters involved in the embodiments of the present application can be stored using cloud storage technology. Cloud storage is a new concept extended and developed from the concept of cloud computing. A distributed cloud storage system refers to a storage system that combines a large number of various types of storage devices (or storage nodes) in the network through functions such as cluster applications, grid technology, and distributed file systems, and collaborates through application software or application interfaces to jointly provide data storage and service access functions to the outside world.
[0096] In a possible application scenario, in order to facilitate reducing the communication delay of retrieval, servers 302 can be deployed in each region, or for load balancing, different servers 302 can respectively serve terminal devices 301 in different regions. For example, the terminal device 301 is located at location a and establishes a communication connection with the server 302 serving location a. The terminal device 301 is located at location b and establishes a communication connection with the server 302 serving location b. Multiple servers 302 form a data sharing system, and data sharing is achieved through blockchain.
[0097] For each server 302 in the data sharing system, it has a node identifier corresponding to that server 302. Each server 302 in the data sharing system can store the node identifiers of other servers 302 in the data sharing system, so that subsequently, according to the node identifiers of other servers 302, the generated blocks can be broadcast to other servers 302 in the data sharing system. A node identifier list can be maintained in each server 302, and the server 302 name and the node identifier are correspondingly stored in this node identifier list. Among them, the node identifier can be an Internet Protocol (IP) address and any other information that can be used to identify the node.
[0098] Of course, the method provided by the embodiments of the present application is not limited to Figure 1 the application scenarios shown, and can also be used in other possible application scenarios, which are not limited by the embodiments of the present application. For Figure 1 the functions that can be achieved by each device in the application scenarios shown will be described together in the subsequent method embodiments, and will not be elaborated here too much.
[0099] The method flows provided in the embodiments of the present application can be executed by the server 302 or the terminal device 301, or can be jointly executed by the server 302 and the terminal device 301.
[0100] Referring to Figure 2 shown, it is a schematic flowchart of a method for identifying regional attributes provided in the embodiments of the present application. This method is applied to a terminal device or a server, and the specific process is as follows:
[0101] S201. Feature extraction is respectively performed on at least two types of modal data extracted from the multimedia resource to obtain corresponding initial features.
[0102] The at least two types of modalities include a text modality and a non-text modality. The non-text modality includes but is not limited to an audio modality, a visual (video or image) modality, etc.
[0103] Correspondingly, the at least two types of modal data include: text modal data and at least one non-text modal data.
[0104] In the multimedia resource, there are initial data corresponding to at least two types of modalities respectively. For example: for the text modality, there are multiple initial text data, including some or all of the following: title, keywords, introduction, and body text; for the audio modality, there is one initial voice data - all the voices in the multimedia resource; for the image modality, there is one initial visual data - all the images in the multimedia resource; for the video modality, there is one initial visual data - the images extracted from the video of the multimedia resource (such as extracting 5 frames of images per second).
[0105] Since non-text modality data often contains text-related information. For example, although audio and text are different forms of language expression, they can convey the same semantic information; the visual content in the visual modality (image modality or video modality) may include text introductions, subtitles, bullet screens and other text information.
[0106] Therefore, by converting the initial non-text data of a non-text modality, corresponding converted text data can be obtained. For example, performing automatic speech recognition (ASR) on the initial audio data to obtain the corresponding converted text data (audio-to-text data); performing optical character recognition (OCR) on the initial visual data of the visual modality to obtain the corresponding converted text data (visual-to-text data).
[0107] For multimedia resources, the extracted text modality data can be the initial text data, or the converted text data, or can include both the initial text data and the converted text data.
[0108] If a multimedia resource has multiple non-text modalities, some of the initial non-text data of the non-text modalities can be selected as the corresponding non-text modality data. For example, for a multimedia resource in video form, which has two non-text modalities, the audio modality and the visual modality, only the initial non-text data of the visual modality (initial visual data) is selected as the corresponding non-text modality data.
[0109] If a multimedia resource has only one non-text modality, the initial non-text data of this one non-text modality can be directly used as the non-text modality data. For example, for a multimedia resource in graphic form, which has one non-text modality, the visual modality, the initial non-text data of the visual modality (initial visual data) is directly selected as the corresponding non-text modality data.
[0110] In some alternative embodiments, when extracting features from each modality data, the modality data can be extracted according to the feature extraction method set for the corresponding modality to obtain the corresponding initial features.
[0111] In practice, taking the text modality data and the visual modality data as examples, first obtain the text modality data corresponding to the multimedia resource and the visual modality data corresponding to the multimedia resource; then extract features from the text modality data according to the feature extraction method set for the text modality to obtain the corresponding initial features, and extract features from the visual modality data according to the feature extraction method set for the visual modality to obtain the corresponding initial features.
[0112] The feature extraction method for the text modality can be implemented using a pre-trained language model based on bidirectional encoding (Bidirectional Encoder Representations from Transformers, BERT), a long short-term memory network (Long Short-Term Memory, LSTM), or a convolutional neural network (Convolutional Neural Networks, CNN), but is not limited thereto.
[0113] Taking BERT as an example, BERT is a pre-trained language model based on the Transformer architecture. Its input data is text modality data, which is segmented into sub-word level sub-data (denoted as tokens), and the output data can be the representation of the sub-data. BERT divides the input text into sentences and adds special tokens at the beginning of each sentence. For example, [CLS] represents the start of a sentence, and [SEP] represents the separation of sentences. In addition, for pre-training, BERT also introduces a Masked Language Model (MLM) task, that is, randomly masking some words in the input text, and then predicting the masked words through the BERT model.
[0114] For the sentence-level representation, in the output of the BERT model, the hidden state of the first word of each sentence is used as the representation of the entire sentence, that is, the sentence-level representation, which can be used for tasks such as text classification and sentiment analysis; for the sub-word level representation, in the output of the BERT model, the hidden state of each sub-word is retained, that is, the sub-word level representation. These representations can be used for sub-word level tasks, such as performing feature fusion in subsequent steps.
[0115] The feature extraction method for the visual modality can be implemented using a visual feature extraction network, such as the Network Vector of Locally Aggregated Descriptors (NetVLAD), the Next-Generation Vector of Locally Aggregated Descriptors (NeXtVLAD), or the Contrastive Language-Image Pre-training (CLIP) model, but is not limited thereto.
[0116] Taking CLIP as an example, CLIP aligns the embedding spaces of images and texts through contrastive learning to achieve feature extraction for visual modality data.
[0117] First, perform contrastive pre-training. Refer to Figure 3 to obtain a large number of image-text pairs. Input the text into the text encoder to obtain the corresponding text encodings (T1 to T N ), and input the image into the image encoder to obtain the corresponding image encodings (I1 to I N ). Through contrastive learning, map the image and text to the same semantic space. The image encoder and text encoder are trained simultaneously with the goal of maximizing the feature similarity of the matching image-text pairs and minimizing the similarity of the non-matching image-text pairs.
[0118] Next, create a dataset classifier. Refer to Figure 4 As shown, the user may define text labels (such as action descriptions) for video tasks and use the trained text encoder to generate label embeddings (t1 to t m ) of these text labels.
[0119] Finally, perform zero-shot prediction. Refer to Figure 5 As shown, input each image in the visual modality data into the trained image encoder in turn, encode each image to obtain the corresponding image encoding (i j ). By comparing each image encoding with each label embedding respectively, obtain the corresponding initial features. Figure 5 Taking the image encoding i1 of the first image as an example, its similarity with the label embedding t3 is the highest, and t3 is used as the initial feature of the first image.
[0120] S202. Perform cross-modal feature fusion on the obtained at least two initial features to obtain fused features.
[0121] Cross-modal feature fusion can be implemented using but not limited to the attention mechanism. The attention mechanism originated from the study of human vision and is a technology widely used in the fields of computer vision (CV) and natural language processing (NLP). It assigns different weights to different parts of the input features, enabling the model to pay more attention to important information, thereby improving the performance and expressive ability of the model.
[0122] In some alternative embodiments, step S202 can be implemented using but not limited to the following methods:
[0123] First, align the dimensions of the at least two initial features.
[0124] Exemplarily, the initial features of each non-text modality are respectively input into the corresponding first fully connected layer (FC layer); the first fully connected layer projects the initial features of a non-text modality into the same dimensional space as the initial features of the text modality (e.g., through a linear layer or MLP) to eliminate the dimensional differences between different modalities.
[0125] Next, for each initial feature after dimension alignment, the following is respectively performed: in an initial feature, modal feature elements and position feature elements are added to obtain the corresponding target feature.
[0126] Among them, the modal feature elements are used to describe the modality of the modal data corresponding to an initial feature; the position feature elements are used to describe the positions of the respective sub-data included in the modal data corresponding to an initial feature in the corresponding modal data.
[0127] Exemplarily, the modal data of each modality contains multiple sub-data. For example: each sub-word level token in the text modal data is used as a sub-data; each segment of speech in the audio modal data is used as a sub-data; each image in the visual modal data is used as a sub-data;
[0128] Feature extraction is performed on the modal data of each modality to obtain the corresponding initial features; each initial feature contains multiple sub-features, and each sub-feature represents a piece of sub-data in the corresponding modal data.
[0129] For each sub-feature, the corresponding modal feature elements are added to enhance the representation of modality-specific information. For example, through parameter learning, the modality-specific prior knowledge (such as the importance of the spatial structure for images) is captured to obtain the modal feature elements. The modal feature elements added to all sub-features in the same initial feature are the same.
[0130] For each sub-feature, position feature elements are added. For example: for a sub-feature of the text modality, its position feature elements describe the position of the corresponding sub-word level token in the sentence; for a sub-feature of the visual modality, its position feature elements describe the order of the corresponding image among all images in the visual modality; for a sub-feature of the audio modality, its position feature elements describe the time order of the corresponding audio segment among all audios. The obtained target feature contains multiple feature elements, and each feature element contains a sub-feature, the corresponding modal feature element, and the position feature element corresponding to this sub-feature.
[0131] Taking the text modality as an example, refer to Figure 6 as shown, the initial feature F of the text modality text contains M sub-features: F1 text ~F M text ; the target feature Ω of the text modality textIt contains M feature elements: Ω1 text ~Ω M text ; among them, for Ω i text , it contains sub-feature F i text , the modal feature element E corresponding to the text modality text and the position feature element T corresponding to this sub-feature i text , i = 1, 2,..., M.
[0132] Taking the visual modality as an example, refer to Figure 7 as shown, the initial feature F of the visual modality visual′ contains M sub-features: F1 visual′ ~F M visua′l ; Align its dimensions with the initial feature of the text modality to obtain the aligned initial feature F visual contains M sub-features: F1 visual ~F M visual ; The target feature Ω of the visual modality visual contains M feature elements: Ω1 visual ~Ω M visual ; among them, for Ω i visual , it contains sub-feature F i visual , the modal feature element E corresponding to the text modality visual and the position feature element T corresponding to this sub-feature i visual , i = 1, 2,..., M.
[0133] Finally, perform cross-modal feature fusion based on the obtained at least two target features to obtain the fused feature.
[0134] In implementation, the process of cross-modal feature fusion can be achieved through the attention mechanism of the Transformer model.
[0135] In some alternative implementation manners, the cross-modal feature fusion can be achieved by but not limited to the following methods:
[0136] Perform splicing based on at least two target features to obtain the first spliced feature;
[0137] For each feature element within the first splicing feature, the following operations are respectively performed: Determine the attention weights between one feature element and each feature element within the first splicing feature, where each attention weight represents the dependency relationship between two feature elements; Based on each feature element and its corresponding attention weight, obtain a fused element corresponding to the feature element.
[0138] Based on the obtained multiple fused elements, obtain the fused feature.
[0139] Taking the cross-modal feature fusion of the text modality and the visual modality through the Transformer model as an example, refer to Figure 8 as shown:
[0140] Before the target feature Ω text (Ω1 text ~Ω M text ) at the beginning of the sequence, add the feature element Ω[CLS] with the CLS marker. Since the Transformer is sensitive to position, by placing Ω[CLS] at the beginning of the sequence, its position is made independent of the task, avoiding the introduction of noise due to position offset, and providing a global semantic representation for the classification task.
[0141] Between the target feature Ω text of the text modality and the target feature Ω visual (Ω1 visual ~Ω M visual ) of the visual modality, add the feature element Ω[unused] with the unused marker for modality separation.
[0142] After the target feature Ω visual , add the feature element Ω[SEP] with the SEP marker to indicate the end of the sequence.
[0143] After the above information splicing, obtain the first splicing feature as the input of the Transformer model.
[0144] Refer to Figure 8 as shown. Through the multi-layer encoder of the Transformer model for cross-modal feature fusion, output the fused elements corresponding to each feature element: the fused element H[CLS] with the CLS marker, the fused elements H1 text ~H M text corresponding to the text modality, the fused elements H1 visual ~H M visual corresponding to the visual modality, the fused element H[unused] with the unused marker, and the fused element H[SEP] with the SEP marker.
[0145] In the embodiments of the present application, each feature element in the first spliced feature is adjusted by an attention weight. The attention weight can adjust the degree of attention of one feature element to the information in other feature elements, learn the dependencies of feature elements within the same modality, and learn the associations of feature elements in different modalities, focusing on the most relevant parts of the feature elements; after multiple layers of interaction, the fusion elements at each position contain global context information, realizing multi-modal deep fusion.
[0146] Based on the obtained multiple fusion elements, a fusion feature is obtained. Taking the fusion elements output by the Transformer model shown above Figure 8 as an example, the fusion feature can be obtained in multiple ways. Here, two ways are taken as examples for illustration:
[0147] Way 1:
[0148] Directly select the fusion element marked by CLS as the fusion feature.
[0149] Since CLS is a special symbol added to the beginning of the input sequence, each layer of the Transformer gradually fuses global information through multi-head attention. As the number of model layers deepens, the representation of CLS will accumulate the context information of the entire sequence and finally form a fusion element that synthesizes the global semantics. Therefore, the marked fusion element can provide a global semantic representation for the classification task, and the fusion element marked by CLS (H[CLS]) can be directly selected as the fusion feature.
[0150] Way 2:
[0151] Based on the fusion elements corresponding to the text modality and the fusion elements corresponding to the visual modality, a fusion feature is calculated.
[0152] Since each of the fusion elements corresponding to the text modality and each of the fusion elements corresponding to the visual modality are obtained by the attention mechanism learning the dependencies of feature elements within the same modality and learning the associations of feature elements in different modalities, focusing on the most relevant parts of the feature elements. Therefore, by synthesizing the fusion elements corresponding to the text modality and the fusion elements corresponding to the visual modality, a fusion feature with a global semantic representation can also be obtained.
[0153] For example, after pooling and splicing the multiple fusion elements corresponding to the text modality and the multiple fusion elements corresponding to the visual modality respectively, a fusion feature is obtained. The formula is expressed as:
[0154] Text fusion sequence htext = AvgPool(H1 text , H2 text , ……H M text );
[0155] The visual fusion sequence hvisual = AvgPool(H1 visual , H2 visual , ……H M visual );
[0156] The fused feature hjoint = [htext; hvisual].
[0157] S203. For the initial features of each non - text modality, perform the following respectively: globally extract features from multiple sub - features in an initial feature to obtain the corresponding non - text global feature.
[0158] As described above, each sub - feature represents: a sub - data item in the corresponding modality data.
[0159] In implementation, since the cross - modal feature fusion process will focus more on the text modality, based on this, this embodiment shows the introduction of non - text global features of non - text modalities.
[0160] Exemplarily, by performing an average pooling operation on the initial features of each non - text modality, globally extract features from multiple sub - features in the initial feature of a non - text modality to obtain the corresponding non - text global feature.
[0161] S204. Based on the fused feature, combined with at least one obtained non - text global feature, obtain the regional recognition result of the multimedia resource.
[0162] By globally extracting features from the initial features of non - text modalities, the expression of the initial features of non - text modalities is strengthened. In this way, based on the fused feature and non - text global features, the accuracy of the regional recognition result is effectively improved.
[0163] In some alternative implementation manners, step S204 can be implemented by, but not limited to, the following methods:
[0164] First, align the dimensions of the fused feature and at least one non - text global feature and then splice them to obtain the second spliced feature.
[0165] Exemplarily, input each non - text global feature into the corresponding second fully - connected layer; the second fully - connected layer projects a non - text global feature into the same dimensional space as the fused feature to eliminate the dimensional differences between these features.
[0166] Then, perform a classification mapping on the second spliced feature to obtain the target classification information.
[0167] Exemplarily, the second splicing feature is input into the third fully-connected layer. The third fully-connected layer performs a linear transformation on the second splicing feature, maps the second splicing feature to the classification space, and outputs the target classification information. It is expressed by the formula:
[0168] The target classification information logits = W·x + b; where W is the weight matrix, b is the bias vector, and x is the second splicing feature.
[0169] Finally, the region recognition result corresponding to the target classification information is used as the region recognition result of the multimedia resource.
[0170] Exemplarily, the target classification information is transformed into the probability distribution of region recognition through the softmax function. The formula is expressed as:
[0171] The probability distribution of region recognition Softmax(z i ) = Softmax(logits); Softmax is the softmax function, and logits is the above-mentioned target classification information;
[0172] The corresponding region recognition result is obtained based on the probability distribution of region recognition.
[0173] Taking the two modalities of text modality and visual modality as examples, refer to Figure 9 as shown:
[0174] The visual global feature Y' visual passes through the second fully-connected layer to align the dimensions with the fusion feature hjoint, and the aligned visual global feature Y is obtained visual ;
[0175] The aligned visual global feature Y visual is spliced with the fusion feature hjoint to obtain the second splicing feature hY = [hjoint; Y visual .
[0176] The second splicing feature hY is input into the third fully-connected layer. The third fully-connected layer outputs the target classification information, and the region recognition result corresponding to the target classification information is used as the region recognition result of the multimedia resource.
[0177] S205. When the region recognition result indicates that the multimedia resource has regionality, based on the text modality data, the target region name associated with the regional event corresponding to the multimedia resource is recognized.
[0178] In implementation, when the multimedia resource has regionality, a large model can be adopted to perform the task of identifying the target regional name associated with the regional event corresponding to the multimedia resource through prompt guidance. The prompt guidance includes information such as the definition of the regional event, the task instruction (how to identify the target regional name associated with the regional event corresponding to the multimedia resource), and the output format of the target regional name (for example, the output target regional name includes the regional word segments corresponding to the respective preset regional hierarchy relationships).
[0179] In some optional implementation manners, step S205 can be implemented by, but not limited to, the following methods:
[0180] Perform associated region analysis of the regional event based on the text modality data to obtain the original regional name, where the original regional name includes at least one original regional word segment;
[0181] Based on the respective original regional levels corresponding to at least one original regional word segment, and in combination with the preset regional hierarchy relationship, perform identification of missing regional levels on the original regional name to obtain the identification result;
[0182] When the identification result indicates that the original regional name does not lack the to-be-filled regional level, use the original regional name as the target regional name;
[0183] When the identification result indicates that the original regional name lacks at least one to-be-filled regional level, from the regional word segment relationship library, for each to-be-filled regional level respectively, obtain the to-be-filled regional word segments associated with at least one original regional word segment, and respectively fill each obtained to-be-filled regional word segment into the original regional name to obtain the target regional name.
[0184] Exemplarily, this embodiment needs to output the target regional name including the regional word segments corresponding to the respective preset regional hierarchy relationships. For example, the regional hierarchy relationship includes four levels: country-province-prefecture-level city-county (county-level cities, counties, and districts are all at the county level), and the target regional name includes the regional word segments corresponding to these four levels respectively.
[0185] Based on this, after performing associated region analysis of the regional event on the text modality data to obtain the original regional name including at least one original regional word segment, in combination with the above regional hierarchy relationship, perform identification of missing regional levels on the original regional name, that is, determine whether the original regional name lacks the to-be-filled regional level, to obtain the identification result. Exemplarily, the to-be-filled regional level is the level lacking in the original regional name, and moreover, the to-be-filled regional level is the level that can be filled based on the original regional word segment, such as the to-be-filled regional level is higher than the level of at least one original regional level.
[0186] During implementation, if the recognition result indicates that the original geographical name does not lack the geographical level to be filled, it means that the original geographical name contains the original geographical word segmentation with a complete hierarchy, or the missing geographical level is a level that cannot be filled based on the original geographical word segmentation. In this case, directly use the original geographical name as the target geographical name. Still taking the geographical hierarchy relationship including the above four levels as an example, if the original geographical name contains the original geographical word segmentation corresponding to each of these four levels and does not lack the geographical level to be filled; if the original geographical name contains the original geographical word segmentation corresponding to the country, province, and prefecture-level city respectively, and the missing level is the lowest-level county, and it is impossible to fill the geographical word segmentation of the lower level through the original geographical word segmentation of the upper level, it is also regarded as not lacking the geographical level to be filled.
[0187] Conversely, if the recognition result indicates that the original geographical name lacks at least one geographical level to be filled, it means that the original geographical name contains incomplete original geographical word segmentation, and the missing level is a geographical level to be filled that can be filled based on the original geographical word segmentation. Therefore, for each geographical level to be filled, in combination with the geographical word segmentation relationship library, obtain the geographical word segmentation to be filled associated with at least one original geographical word segmentation (such as the original geographical word segmentation corresponding to the original geographical level lower than the geographical level to be filled), and fill the geographical word segmentation to be filled into the original geographical name to obtain a target geographical name with a more complete hierarchy than the original geographical name. Still taking the geographical hierarchy relationship including the above four levels as an example, if the original geographical name contains the original geographical word segmentation corresponding to the country, province, and county respectively, and the missing level is the prefecture-level city, use the original geographical word segmentation of the lower level (the original geographical word segmentation corresponding to the county) to fill the geographical word segmentation corresponding to the prefecture-level city.
[0188] In some optional implementation manners, the original geographical name can be obtained through, but not limited to, the following methods:
[0189] When the text modal data contains one item of data, perform an associated geographical analysis of geographical events on the one item of data to obtain the original geographical name; the one item of data is: an initial text data item, or a transformed text data corresponding to a non-text modality.
[0190] During implementation, if the text modal data has only one item of data, directly perform an associated geographical analysis of geographical events on it to obtain the original geographical name.
[0191] When the text modal data contains multiple items of data, based on the respective priorities of the multiple items of data, sequentially perform an associated geographical analysis of geographical events on the multiple items of data to obtain the original geographical name; the multiple items of data are any one of the following: multiple initial text data items; transformed text data corresponding to multiple non-text modalities; at least one initial text data item and transformed text data corresponding to at least one non-text modality respectively.
[0192] In implementation, if the text modality data contains multiple pieces of data, the importance of these multiple pieces of data for geographical location recognition is different. Therefore, priorities are set for different data types, and based on the priorities, the original geographical location names can be obtained more efficiently and accurately.
[0193] Exemplarily, the priorities of the initial text data are: title > body text; the priorities of the converted text data are audio-to-text data > visual-to-text data; the priorities of all initial text data are greater than the priorities of any converted text data.
[0194] In some alternative implementation manners, for the case where the text modality data contains multiple pieces of data, to obtain the original geographical location name, it can be implemented by, but not limited to, the following methods:
[0195] Read multiple pieces of data in sequence according to the priority order. Among them, for each piece of data read, perform the following operations:
[0196] First, perform an associated geographical location analysis on a piece of data for geographical location events to obtain at least one geographical location word segment.
[0197] Next, when a piece of data is the first piece of data read, use at least one geographical location word segment as the geographical location name corresponding to the piece of data; when a piece of data is not the first piece of data read, supplement the geographical location name corresponding to the previously read piece of data based on at least one geographical location word segment to obtain the geographical location name corresponding to the piece of data.
[0198] For example: Compare the first levels of each of at least one geographical location word segment with the second levels of the geographical location word segments included in the geographical location name corresponding to the previous piece of data. If there are new first levels compared to the second levels, add the geographical location word segments corresponding to the new first levels to the geographical location name corresponding to the previous piece of data to obtain the geographical location name corresponding to this piece of data.
[0199] Finally, when the geographical location name corresponding to a piece of data contains geographical location word segments at the target level in the geographical location hierarchy, or when a piece of data is the last piece of data read, obtain the original geographical location name based on the geographical location name corresponding to the piece of data.
[0200] In implementation, if the geographical location name corresponding to this piece of data already contains geographical location word segments at the target level (such as the lowest level in the geographical location hierarchy), then the geographical location name corresponding to this piece of data may already contain geographical location word segments of the complete hierarchy; it may also contain geographical location word segments of an incomplete hierarchy, and the missing part is a geographical location level that can be filled based on the geographical location name corresponding to this piece of data.
[0201] Based on this, after the geographical word segmentation of the geographical name corresponding to a piece of data contains the target level in the geographical hierarchy relationship, instead of reading the subsequent data, the original geographical name is obtained based on the geographical name corresponding to this piece of data.
[0202] During implementation, if the last piece of data has been read, regardless of the level of the geographical word segmentation contained in the geographical name corresponding to this piece of data, the original geographical name is obtained based on the geographical name corresponding to this piece of data.
[0203] For example, taking the geographical name corresponding to this piece of data as the original geographical name, that is to say, each geographical word segmentation in the geographical name corresponding to this piece of data is used as the original geographical word segmentation.
[0204] Taking news A in the form of a video as an example, the geographical attribute recognition process can be referred to Figure 10 as shown:
[0205] The news title 3 of news A is "The Changing Times of the White Pagoda", and there are buildings named White Pagoda in multiple cities, that is, there is a situation of duplicate building names;
[0206] The video of news A contains the historical evolution of the White Pagoda and an introduction to the city where the White Pagoda is located.
[0207] Perform feature extraction corresponding to the text modality on the news title 3, the text obtained by converting the audio, and the text obtained by converting the video (including text introductions, subtitles, and bullet screens) to obtain the initial feature F of the text modality text (including M sub-features: F1 text ~F M text );
[0208] Sample the video of news A, extract M frames of images, perform feature extraction on the M frames of images, and obtain the initial feature F of the visual modality visual′ (including M sub-features: F1 visual ~F M visual ); Align its dimensions with the initial feature of the text modality to obtain the aligned initial feature F visual including M sub-features: F1 visual ~F M visual ;
[0209] In the initial feature F visual add the modality feature element E corresponding to the text modality text and the position feature element T corresponding to each sub-feature i text , to obtain the target feature Ω text , including M feature elements: Ω1text ~ΩM text 。
[0210] In the initial feature F visual add the modality feature element E corresponding to the visual modality visual and the position feature element T corresponding to each sub-feature i visual to obtain the target feature Ω visual which contains M feature elements: Ω1 visual ~Ω M visual 。
[0211] Concatenate the target feature Ω text with the target feature Ω visual to obtain the first concatenated feature [Ω text ; Ω visual .
[0212] Perform cross-modal feature fusion on the first concatenated feature to obtain the fused feature hjoint.
[0213] Extract the global feature from the initial feature F visual to obtain the visual global feature Y ′visual 。
[0214] Align the dimension of the visual global feature Y ′visual with the fused feature hjoint to obtain the aligned visual global feature Y visual 。
[0215] Concatenate the aligned visual global feature Y visual with the fused feature hjoint to obtain the second concatenated feature hY = [hjoint; Y visual .
[0216] Perform classification mapping on the second concatenated feature to obtain the target classification information.
[0217] The regional recognition result corresponding to the target classification information is that it has regionality, that is, news A has regionality.
[0218] Based on the above text modality data, conduct an associated regional analysis of regional events to obtain the original regional name, which includes two original regional participles: Country AA and County DD.
[0219] The preset regional hierarchy relationship includes country-province-prefecture-level city-county. Identify the missing regional levels of the original regional name, and determine that the original regional name lacks the two regional levels of province and prefecture-level city to be filled;
[0220] From the regional word segmentation relationship database, obtain the to-be-filled regional word segmentations - Province BB and City CC associated with County DD, and fill these two to-be-filled regional word segmentations into the original regional name to obtain a target regional name containing the four regional word segmentations: Country AA, Province BB, City CC, and County DD.
[0221] Based on the same inventive concept, an embodiment of the present application provides a regional attribute recognition device. As Figure 11 shown, it is a schematic structural diagram of the regional attribute recognition device 1100, which may include:
[0222] A feature extraction module 1101, configured to perform feature extraction on at least two types of modal data extracted based on multimedia resources respectively to obtain corresponding initial features; the at least two types of modal data include: text modal data and at least one non-text modal data; the text modal data includes at least one of the following: at least one initial text data in the multimedia resources, and transformed text data obtained by modal transformation of at least one non-text initial non-text data in the multimedia resources;
[0223] A feature fusion module 1102, configured to perform cross-modal feature fusion on the obtained at least two initial features to obtain a fusion feature;
[0224] A non-text processing module 1103, configured to respectively perform, for the initial feature of each non-text modality: globally extract features from multiple sub-features in one initial feature to obtain corresponding non-text global features; each sub-feature represents: one sub-data in the corresponding non-text modal data;
[0225] A regionality recognition module 1104, configured to obtain a regional recognition result of the multimedia resource based on the fusion feature and in combination with the obtained at least one non-text global feature;
[0226] A regional name recognition module 1105, configured to, when the regional recognition result indicates that the multimedia resource has regionality, recognize a target regional name associated with the regional event corresponding to the multimedia resource based on the text modal data.
[0227] In some optional embodiments, the feature fusion module 1102 is specifically configured to:
[0228] Align the dimensions of the at least two initial features;
[0229] For each initial feature after dimension alignment, the following operations are respectively performed: In an initial feature, modal feature elements and position feature elements are added to obtain a corresponding target feature; the modal feature elements are used to describe the modality of the modal data corresponding to the one initial feature; the position feature elements are used to describe the positions of the respective sub-data included in the modal data corresponding to the one initial feature in the modal data.
[0230] Based on at least two obtained target features, cross-modal feature fusion is performed to obtain the fusion feature.
[0231] In some optional implementation manners, the feature fusion module 1102 is specifically configured to:
[0232] Based on the at least two target features, splicing is performed to obtain a first spliced feature;
[0233] For each feature element in the first spliced feature, the following operations are respectively performed: Attention weights between a feature element and each feature element in the first spliced feature are respectively determined, and each attention weight represents the dependency relationship between two feature elements; based on the respective feature elements and the attention weights corresponding to the respective feature elements, a fusion element corresponding to the one feature element is obtained;
[0234] Based on the obtained multiple fusion elements, the fusion feature is obtained.
[0235] In some optional implementation manners, the regional identification module 1104 is specifically configured to:
[0236] The fusion feature and the at least one non-text global feature are dimension-aligned and then spliced to obtain a second spliced feature;
[0237] Classification mapping is performed on the second spliced feature to obtain target classification information;
[0238] The regional identification result corresponding to the target classification information is used as the regional identification result of the multimedia resource.
[0239] In some optional implementation manners, the regional name identification module 1105 is specifically configured to:
[0240] Based on the text modal data, associated region analysis of regional events is performed to obtain an original regional name, and the original regional name includes at least one original regional word segment;
[0241] Based on the original regional levels corresponding to the at least one original regional word segment and in combination with a preset regional hierarchy relationship, regional level missing identification is performed on the original regional name to obtain an identification result;
[0242] When the recognition result indicates that the original geographical name does not lack the geographical level to be filled, use the original geographical name as the target geographical name;
[0243] When the recognition result indicates that the original geographical name lacks at least one geographical level to be filled, from the geographical word segmentation relationship library, for each of the geographical levels to be filled, obtain the geographical word segmentation to be filled associated with the at least one original geographical word segmentation, and respectively fill each obtained geographical word segmentation to be filled into the original geographical name to obtain the target geographical name.
[0244] In some alternative embodiments, the geographical name recognition module 1105 is specifically configured to:
[0245] When the text modal data contains one item of data, perform an associated geographical analysis of geographical events on the one item of data to obtain the original geographical name; the one item of data is: one item of initial text data, or, a transformed text data corresponding to a non-text modality;
[0246] When the text modal data contains multiple items of data, based on the respective priorities of the multiple items of data, sequentially perform an associated geographical analysis of geographical events on the multiple items of data to obtain the original geographical name; the multiple items of data are any one of the following: multiple items of initial text data; transformed text data respectively corresponding to multiple non-text modalities; at least one item of initial text data and transformed text data respectively corresponding to at least one non-text modality.
[0247] In some alternative embodiments, the geographical name recognition module 1105 is specifically configured to:
[0248] Read the multiple items of data in order of priority. Among them, each time one item of data is read, perform the following operations:
[0249] Perform an associated geographical analysis of geographical events on the one item of data to obtain at least one geographical word segmentation;
[0250] When the one item of data is the first read data, use the at least one geographical word segmentation as the geographical name corresponding to the one item of data;
[0251] When the one item of data is not the first read data, supplement the geographical name corresponding to the previous read data based on the at least one geographical word segmentation to obtain the geographical name corresponding to the one item of data;
[0252] When the geographical name corresponding to the one item of data contains the geographical word segmentation of the target level in the geographical hierarchy relationship, or, the one item of data is the last read data, obtain the original geographical name based on the geographical name corresponding to the one item of data.
[0253] For the convenience of description, the above - mentioned parts are divided into respective modules (or units) according to their functions and described separately. Of course, when implementing this application, the functions of the respective modules (or units) can be implemented in the same or multiple software or hardware.
[0254] Regarding the device in the above - mentioned embodiments, the specific manner in which each module executes the request has been described in detail in the embodiments related to the method, and will not be elaborated here.
[0255] Based on the same technical concept, an embodiment of this application provides a computer device, which can be Figure 1 the terminal device and / or server shown, such as Figure 12 shown, including at least one processor 1201 and a memory 1202 connected to at least one processor. In the embodiments of this application, the specific connection medium between the processor 1201 and the memory 1202 is not limited. Figure 12 Taking the example that the processor 1201 and the memory 1202 are connected through a bus. The bus can be divided into an address bus, a data bus, a control bus, etc.
[0256] In the embodiments of this application, the memory 1202 stores instructions executable by at least one processor 1201. By executing the instructions stored in the memory 1202, at least one processor 1201 can execute the steps of the above - mentioned geographical attribute recognition method.
[0257] Among them, the processor 1201 is the control center of the computer device. It can connect various parts of the computer device through various interfaces and lines. By running or executing the instructions stored in the memory 1202 and calling the data stored in the memory 1202, the geographical attribute recognition can be achieved. Optionally, the processor 1201 may include one or more processing units. The processor 1201 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, the object interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above - mentioned modem processor may not be integrated into the processor 1201. In some embodiments, the processor 1201 and the memory 1202 can be implemented on the same chip, and in some embodiments, they can also be separately implemented on independent chips.
[0258] The processor 1201 may be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application specific integrated circuit (ASIC), a field programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.
[0259] As a non-volatile computer-readable storage medium, the memory 1202 can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The memory 1202 may include at least one type of storage medium, for example, it may include flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic memory, magnetic disk, optical disk, and so on. The memory 1202 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer device, but is not limited thereto. The memory 1202 in the embodiments of the present application may also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data.
[0260] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program executable by a computer device. When the program runs on the computer device, the computer device is caused to execute the steps of the above-mentioned geographical attribute recognition method.
[0261] Based on the same inventive concept, an embodiment of the present application provides a computer program product. The computer program product includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer device, the computer device is caused to execute the steps of the above-mentioned geographical attribute recognition method.
[0262] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0263] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer device or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0264] These computer program instructions can also be stored in a computer-readable memory that can direct a computer device or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means realizes the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0265] These computer program instructions can also be loaded onto a computer device or other programmable data processing devices, so that a series of operation steps are executed on the computer device or other programmable devices to generate a process implemented by the computer device. Thus, the instructions executed on the computer device or other programmable devices provide steps for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0266] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the present invention.
[0267] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A method for identifying geographical attributes, characterized in that, Including: Performing feature extraction on at least two modalities of data extracted based on multimedia resources respectively to obtain corresponding initial features; The at least two modalities of data include: text modality data and at least one non-text modality data; the text modality data includes at least one of the following: at least one initial text data in the multimedia resource, and transformed text data obtained by transforming at least one initial non-text data of a non-text modality in the multimedia resource; Performing cross-modal feature fusion on the obtained at least two initial features to obtain a fused feature; For the initial feature of each non-text modality, respectively perform: globally extracting features from multiple sub-features in an initial feature to obtain corresponding non-text global features; each sub-feature represents: a sub-data item in the corresponding non-text modality data; Based on the fused feature, combining the obtained at least one non-text global feature to obtain a geographical location recognition result of the multimedia resource; When the geographical location recognition result indicates that the multimedia resource has geographical characteristics, based on the text modality data, identifying a target geographical location name associated with a geographical event corresponding to the multimedia resource.
2. The method according to claim 1, characterized in that, Performing cross-modal feature fusion on the obtained at least two initial features to obtain a fused feature, including: Aligning the dimensions of the at least two initial features; For each initial feature after dimension alignment, respectively perform: adding a modality feature element and a position feature element to an initial feature to obtain a corresponding target feature; the modality feature element is used to describe the modality of the modality data corresponding to the initial feature; the position feature element is used to describe the positions of respective sub-data items included in the modality data corresponding to the initial feature in the modality data; Performing cross-modal feature fusion based on the obtained at least two target features to obtain the fused feature.
3. The method according to claim 2, characterized in that, Performing cross-modal feature fusion based on the obtained at least two target features to obtain the fused feature, including: Performing splicing based on the at least two target features to obtain a first spliced feature; For each feature element in the first spliced feature, respectively perform: determining the attention weights between a feature element and each feature element in the first spliced feature, each attention weight representing the dependence relationship between two feature elements; based on the feature elements and the attention weights corresponding to the respective feature elements, obtaining a fused element corresponding to the feature element; Based on the obtained multiple fused elements, obtaining the fused feature.
4. The method according to any one of claims 1 to 3, characterized in that, Based on the fused feature, combining the obtained at least one non-text global feature to obtain a geographical location recognition result of the multimedia resource, including: Aligning the dimensions of the fused feature and the at least one non-text global feature and then performing splicing to obtain a second spliced feature; Performing classification mapping on the second spliced feature to obtain target classification information; Taking the geographical location recognition result corresponding to the target classification information as the geographical location recognition result of the multimedia resource.
5. The method according to any one of claims 1-3, characterized in that, Based on the text modality data, identifying a target geographical location name associated with a geographical event corresponding to the multimedia resource, including: Performing associated region analysis of regional events based on the text modality data to obtain an original region name, where the original region name includes at least one original region word segment; Based on the original region levels corresponding to the at least one original region word segment respectively, and in combination with a preset regional hierarchy relationship, performing missing region level identification on the original region name to obtain an identification result; When the identification result indicates that the original region name does not lack the region level to be filled, using the original region name as the target region name; When the identification result indicates that the original region name lacks at least one region level to be filled, from the region word segment relationship library, for each of the region levels to be filled respectively, obtaining the region word segments to be filled associated with the at least one original region word segment, and respectively filling each obtained region word segment to be filled into the original region name to obtain the target region name.
6. The method according to claim 5, characterized in that, Performing associated region analysis of regional events based on the text modality data to obtain an original region name, including: When the text modality data includes one item of data, performing associated region analysis of regional events on the one item of data to obtain the original region name; the one item of data is: an initial text data item, or a converted text data corresponding to a non - text modality; When the text modality data includes multiple items of data, based on the priorities of the multiple items of data, sequentially performing associated region analysis of regional events on the multiple items of data to obtain the original region name; the multiple items of data are any one of the following: multiple initial text data items; converted text data corresponding to multiple non - text modalities respectively; at least one initial text data item and converted text data corresponding to at least one non - text modality respectively.
7. The method according to claim 6, characterized in that, Sequentially performing associated region analysis of regional events on the multiple items of data to obtain the original region name, including: Reading the multiple items of data in order of priority, where for each item of data read, the following operations are performed: Performing associated region analysis of regional events on the one item of data to obtain at least one region word segment; When the one item of data is the first item of data read, using the at least one region word segment as the region name corresponding to the one item of data; When the one item of data is not the first item of data read, supplementing the region name corresponding to the previous item of data read based on the at least one region word segment to obtain the region name corresponding to the one item of data; When the region name corresponding to the one item of data includes a region word segment of the target level in the regional hierarchy relationship, or the one item of data is the last item of data read, obtaining the original region name based on the region name corresponding to the one item of data.
8. A regional attribute recognition device, characterized in that, Including: A feature extraction module, configured to perform feature extraction on at least two types of modality data extracted from multimedia resources respectively to obtain corresponding initial features; The at least two types of modality data include: text modality data and at least one non-text modality data; the text modality data includes at least one of the following: at least one initial text data in the multimedia resource, and transformed text data obtained by modality transformation of at least one initial non-text data of a non-text modality in the multimedia resource; A feature fusion module, configured to perform cross-modal feature fusion on at least two obtained initial features to obtain fused features; A non-text processing module, configured to, for each initial feature of a non-text modality, respectively perform: globally extracting features from multiple sub-features in an initial feature to obtain corresponding non-text global features; each sub-feature represents: a piece of sub-data in the corresponding non-text modality data; A regional identification module, configured to obtain a regional identification result of the multimedia resource based on the fused features and in combination with at least one obtained non-text global feature; A regional name identification module, configured to, when the regional identification result indicates that the multimedia resource has a region, identify a target regional name associated with a regional event corresponding to the multimedia resource based on the text modality data.
9. A computer device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps of any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that It includes a computer program, and when the computer program runs on a computer device, the computer program is used to cause the computer device to execute the steps of any one of claims 1 to 7.
11. A computer program product, characterized in that, It includes a computer program, the computer program is stored in a computer-readable storage medium, and a processor of a computer device reads and executes the computer program from the computer-readable storage medium, so that the computer device executes the steps of any one of claims 1 to 7.
Citation Information
Cited By
Dialect speech recognition method, storage medium and electronic device
CN121393420A