Event aggregation method, device and equipment based on multimodal data
By classifying events and determining address information of multimodal data, the event aggregation problem of multimodal data is solved, and efficient data aggregation and display are achieved.
Patent Information
- Application Number
- CN202310244719.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-14
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2043-03-14
AI Technical Summary
In a multimodal data scenario, how to effectively aggregate multimodal data events to facilitate the display of various modal data related to the event.
By classifying multimodal data, the respective event classification results and event address information are determined, and clustering is performed based on these results to form event clusters. The event classification results and address information of the modal data in the same cluster are the same.
It achieves efficient aggregation of multimodal data, facilitates the subsequent visualization of various modal data of the same event, and improves the accuracy of the display of event-related data.
Smart Images

Figure CN116310682B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence, specifically natural language processing, deep learning technology, etc., and can be applied in smart city and smart government scenarios, especially to event aggregation methods, devices and equipment based on multimodal data. Background Art
[0002] To provide users with a comprehensive understanding of an event, various modal data related to the event can be displayed to them, such as text information, images, and video data related to the event. When there is multimodal data to be processed, how to aggregate the multimodal data into events is crucial for displaying the various modal data related to the event. Summary of the Invention
[0003] The present disclosure provides a method, apparatus, and device for event aggregation based on multimodal data.
[0004] According to one aspect of the present disclosure, a method for event aggregation based on multimodal data is provided, comprising: acquiring multimodal data to be processed; performing event classification on the multimodal data respectively to obtain event classification results corresponding to each of the multimodal data; determining event address information corresponding to each of the multimodal data; and clustering the multimodal data according to the event classification results and the event address information to obtain at least one clustering cluster, wherein the event classification results corresponding to various modal data in the same clustering cluster and the event address information are the same.
[0005] According to another aspect of the present disclosure, an event aggregation device based on multimodal data is provided, including: an acquisition module for acquiring multimodal data to be processed; an event classification module for performing event classification on the multimodal data respectively to obtain event classification results corresponding to each of the multimodal data; a first determination module for determining event address information corresponding to each of the multimodal data; and a first clustering module for clustering the multimodal data according to the event classification results and the event address information to obtain at least one cluster cluster, wherein the event classification results corresponding to various modal data in the same cluster cluster and the event address information are the same.
[0006] According to another aspect of the present disclosure, a training device for a question-answer matching model is provided, comprising: obtaining training data generated by the event aggregation method based on multimodal data as described above; and training the question-answer matching model based on the training data.
[0007] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the event aggregation method based on multimodal data of the present disclosure.
[0008] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the event aggregation method based on multimodal data disclosed in an embodiment of the present disclosure.
[0009] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, the method for event aggregation based on multimodal data of the present disclosure is implemented.
[0010] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0012] Figure 1 is a schematic diagram according to a first embodiment of the present disclosure;
[0013] Figure 2 is a schematic diagram according to a second embodiment of the present disclosure;
[0014] Figure 3 is a schematic diagram according to a third embodiment of the present disclosure;
[0015] Figure 4 is a schematic diagram according to a fourth embodiment of the present disclosure;
[0016] Figure 5 is a schematic diagram according to a fifth embodiment of the present disclosure;
[0017] Figure 6 It is a block diagram of an electronic device used to implement the event aggregation method based on multimodal data according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0018] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0019] The following describes the event aggregation method, device and apparatus based on multimodal data according to embodiments of the present disclosure with reference to the accompanying drawings.
[0020] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure.
[0021] like Figure 1 As shown, the event aggregation method based on multimodal data may include:
[0022] Step 101: Acquire multimodal data to be processed.
[0023] It should be noted that the executor of the above-mentioned event aggregation method based on multimodal data is an event aggregation device based on multimodal data. The event aggregation device based on multimodal data can be implemented by software and / or hardware. The event aggregation device based on multimodal data in this embodiment can be an electronic device, or can be configured in an electronic device.
[0024] In this example embodiment, electronic devices may include but are not limited to terminal devices, servers and other devices, and this embodiment does not limit the electronic devices.
[0025] The multimodal data to be processed refers to various modal data to be aggregated for events.
[0026] In some examples, the multimodal data to be processed may be multimodal data to be processed in a specified field, that is, the fields to which the multimodal data to be processed belong are the same, and are all specified fields. For example, the specified fields may be various fields such as government affairs, social fields, and medical fields. This embodiment does not specifically limit this.
[0027] Among them, multimodal data can include text modal data, visual modal data, and voice modal data, etc.
[0028] In this example, multimodal data including text modal data and visual modal data is taken as an example for exemplary description.
[0029] In this example, the visual modality data may include image modality data and video modality data.
[0030] Step 102 : performing event classification on the multimodal data respectively to obtain event classification results corresponding to each of the multimodal data.
[0031] In some exemplary embodiments, for each modal data in the multimodal data, the modal data may be subjected to event classification using an event classification algorithm corresponding to the modal data to obtain an event classification result corresponding to the modal data.
[0032] Among them, for each type of modal data, the number of each type of modal data in this example can be greater than or equal to 1. For example, in the case where the multimodal data includes text-type modal data, there can be 10 text-type modal data, that is, there are 10 text modal data to be processed. Correspondingly, an event classification algorithm that can process text modal data can be obtained, and event classification is performed on these 10 text modal data to be processed respectively to obtain event classification results corresponding to each text modal data.
[0033] Step 103: Determine event address information corresponding to each of the multimodal data.
[0034] The event address information represents the address information where the event described by the corresponding modal data occurs.
[0035] Among them, it can be understood that different methods can be used to determine the event address information corresponding to the corresponding modal data for different types of modal data. For example, for text modal data in multimodal data, event address extraction can be performed on the text modal data to obtain the event address information of the text modal data. For visual modal data in multimodal data, the position information of the camera that can be bound to the visual modal data can be determined to determine the address information of the location where the event described by the visual modal data occurs, that is, the event address information corresponding to the visual modal data can be determined. For another example, for voice modal data in multimodal data, the voice modal data can be converted into text to obtain text information corresponding to the voice modal data, and event address extraction can be performed on the text information to determine the event address information corresponding to the voice modal data.
[0036] Step 104 : clustering the multimodal data according to the event classification results and the event address information to obtain at least one cluster, wherein the event classification results and event address information corresponding to the various modal data in the same cluster are the same.
[0037] In this example, by clustering based on event classification results and event address information, various modal data with the same event classification results and the same event address information can be aggregated into one cluster, which facilitates the subsequent visualization of various modal data of the same event in the cluster.
[0038] The event aggregation method based on multimodal data provided by the embodiments of the present disclosure, when processing the multimodal data to be processed, determines the event classification results and event address information corresponding to each of the multimodal data, and clusters the multimodal data to obtain at least one cluster based on the event classification results and event address information. Thus, a method for performing event aggregation on multimodal data based on the event classification results and event address information is provided, conveniently implementing event aggregation on multimodal data.
[0039] In some exemplary embodiments, in order to clearly understand how to classify multimodal data into events to obtain the event classification results corresponding to each of the multimodal data, the following is combined with Figure 2 This process is described exemplarily.
[0040] Figure 2 is a schematic diagram according to the second embodiment of the present disclosure.
[0041] like Figure 2 As shown, the event aggregation method based on multimodal data may include:
[0042] Step 201: Acquire multimodal data to be processed.
[0043] It should be noted that, for the specific implementation of step 201, please refer to the relevant description of the embodiment of the present disclosure, which will not be repeated here.
[0044] Step 202 : For each modal data in the multimodal data, the modal data is input into an event classification model corresponding to the modal data to obtain an event classification result of the modal data.
[0045] In other words, for each type of modal data, an event classification model that can process that type of modal data can be determined, and then the event classification model is used to classify the modal data to obtain an event classification result for that type of modal data. Thus, the event classification model can be used to quickly and accurately determine the event classification result for the corresponding modal data.
[0046] It can be understood that in this example, the event classification models corresponding to different types of modal data are different.
[0047] Among them, the event classification models corresponding to the above different types of modal data are pre-trained.
[0048] Among them, in this example, the training processes of event classification models corresponding to different types of modal data are independent of each other, that is, the event classification models corresponding to different types of modal data can be modeled independently, and the modeling method of each modality is not affected by the models of other modalities.
[0049] As an example, for text modal data, training data can be used to train an initial classification model to obtain a trained event classification model. The training data includes text samples and event classification labels corresponding to the text samples. The initial classification model in this example can be an Enhanced Representation from kNowledge IntEgration (ERNIE) model.
[0050] As an example, for visual modality data, training data can be used to train an initial classification model to obtain a trained event classification model, where the training data can include visual sample data and corresponding event classification labels. In some examples, the initial classification model in this example can be a residual network model or other type of network model, which is not specifically limited in this embodiment.
[0051] Among them, it should be noted that, compared with the technical solution of using the same pre-trained model to understand the content of multimodal data to obtain the content understanding feature vector corresponding to each modal data, and aggregating events for multimodal data based on the similarity between the content understanding feature vectors, the model that can understand the content of multimodal data in this technical solution also needs to be pre-trained. In this example, the event classification model corresponding to each modal data is trained independently. Compared with the solution of using multimodal data to train the model for content understanding of multimodal data, the number of annotations for labeling sample data can be reduced.
[0052] Step 203: Determine the event address information corresponding to each of the multimodal data.
[0053] In some exemplary embodiments, multimodal data includes textual modal data and visual modal data. One possible implementation method for determining the event address information corresponding to each of the multimodal data is to: extract the event address from the textual modal data using a feature extraction model to obtain the event address information corresponding to the textual modal data; and determine the event address information for the visual modal data based on the position information of the camera associated with the visual modal data. Thus, the feature extraction model can quickly and accurately determine the event address information corresponding to the textual modal data, and based on the position information of the camera associated with the visual modal data, the event address information for the visual modal data can be accurately determined.
[0054] Step 204 : clustering the multimodal data according to the event classification results and the event address information to obtain at least one cluster, wherein the event classification results and event address information corresponding to the various modal data in the same cluster are the same.
[0055] It should be noted that, for the specific description of step 204, reference can be made to the relevant description of the embodiment of the present disclosure, which will not be repeated here.
[0056] In this example, when processing multimodal data, for each type of modal data in the multimodal data, an event classification model for processing that type of modal data can be determined, and then the event classification model is used to classify the modal data to obtain an event classification result for that type of modal data. Thus, the event classification model can be used to quickly and accurately determine the event classification result for the corresponding modal data.
[0057] Based on any one of the above embodiments, in some exemplary implementations, when the classification label systems based on which events are classified for various modal data are different, in order to facilitate the subsequent accurate clustering of multimodal data according to the event classification results and event address information, the multimodal data is clustered according to the event classification results and event address information to obtain at least one cluster cluster. The method may also include: determining the original classification label system on which the event classification results are based; and mapping the event classification results according to the mapping relationship between the original classification label system and the preset target classification label system.
[0058] Among them, the target classification label system is a classification label system pre-set in the event aggregation device based on multimodal data.
[0059] Specifically, after determining the original classification label system on which the event classification result is based, it can be determined whether the original classification label system and the target classification label system are the same. If they are not the same, the event classification result is mapped according to the mapping relationship between the original classification label system and the preset target classification label system to obtain the mapped event classification result.
[0060] Correspondingly, the multimodal data may be clustered according to the event address information and the mapped event classification results to obtain at least one cluster.
[0061] In some exemplary embodiments, the related art uses the same pre-trained model to perform content understanding on multimodal data to obtain the content understanding feature vector corresponding to each modal data, and based on the similarity between the content understanding feature vectors, the multimodal data is subjected to event aggregation to obtain event aggregation results. However, how to name the obtained event aggregation results is also a difficult problem to solve. In this example, for each cluster, the cluster name of the cluster is generated based on the event classification results and event address information corresponding to the various modal data in the cluster. Thus, the cluster is named based on the event classification results and event address information.
[0062] As an example, the event classification result and the event address information may be concatenated, and the concatenated result may be used as the cluster name of the cluster.
[0063] Based on any of the above embodiments, in some examples, in order to further aggregate various modal data of the same event, for the cluster cluster, according to the event subject corresponding to the various modal data in the cluster cluster, the various modal data in the cluster cluster are clustered again to obtain at least one cluster sub-cluster corresponding to the cluster cluster, wherein the event subject corresponding to the various modal data in the same cluster sub-cluster is the same. Thus, based on the event subject, the cluster clusters with the same event classification results and event address information can be clustered again, so that the modal data with the same event classification results, event address information and event subject can be aggregated in a cluster sub-cluster, so as to facilitate the subsequent display of event-related data based on the content of the cluster sub-cluster, which can further improve the accuracy of the display of event-related data.
[0064] It is understandable that the event subjects corresponding to each multimodal data can be determined in a variety of ways, as illustrated by the following examples:
[0065] As an example, the event subject corresponding to the corresponding modal data can be obtained based on the pre-saved correspondence between the modal data and the event subject.
[0066] As another example, for text modal data in multimodal data, elements can be extracted from the text modal data through a feature extraction model to obtain feature extraction results of the text modal data, and the event subject of the text modal data can be determined from the feature extraction results.
[0067] The element extraction result may include but is not limited to information such as the event subject, event address information, and event time information, and this embodiment does not specifically limit this.
[0068] For the visual modality data in the multimodal data, event analysis may be performed on the visual modality data to obtain an event analysis result, and the event subject of the visual modality data may be determined based on the event analysis result.
[0069] In order to clearly understand the present disclosure, Figure 3 The event aggregation method based on multimodal data of this embodiment is further exemplarily described.
[0070] Figure 3 is a schematic diagram according to a third embodiment of the present disclosure.
[0071] like Figure 3 As shown, the event aggregation method based on multimodal data may include:
[0072] Step 301: Acquire multimodal data to be processed.
[0073] Step 302 : For each modal data in the multimodal data, the modal data is input into an event classification model corresponding to the modal data to obtain an event classification result of the modal data.
[0074] It should be noted that, for the specific implementation of step 301 and step 302, please refer to the relevant description of the embodiment of the present disclosure, which will not be repeated here.
[0075] Step 303: Determine the original classification label system on which the event classification result is based.
[0076] In some examples, the classification label system based on which the event classification model performs event classification can be used as the original classification label system based on which the event classification result is based.
[0077] Step 304 : Mapping the event classification results according to the mapping relationship between the original classification label system and the preset target classification label system to obtain a mapped event classification result.
[0078] Step 305: Determine the event address information corresponding to each of the multimodal data.
[0079] For the specific implementation of step 305, please refer to the relevant description in the embodiment of the present disclosure, which will not be repeated here.
[0080] Step 306 : Cluster the multimodal data according to the event address information and the mapped event classification results to obtain at least one cluster, wherein the event classification results and event address information corresponding to various modal data in the same cluster are the same.
[0081] Step 307 : For the cluster, generate a cluster name for the cluster according to the event classification results and event address information corresponding to various modal data in the cluster.
[0082] In this example, a method for event aggregation of multimodal data is provided, and after multimodal data with the same event address information and event classification results are aggregated into a cluster, the cluster is named according to the event address information and event classification results.
[0083] In order to implement the above embodiment, the embodiment of the present disclosure also provides an event aggregation device based on multimodal data.
[0084] Figure 4 is a schematic diagram according to a fourth embodiment of the present disclosure.
[0085] like Figure 4 As shown, the event aggregation device 400 based on multimodal data may include: an acquisition module 401, an event classification module 402, a first determination module 403 and a first clustering module 404, wherein:
[0086] An acquisition module 401 is used to acquire multimodal data to be processed;
[0087] An event classification module 402 is used to classify the multimodal data into events to obtain event classification results corresponding to the multimodal data.
[0088] A first determining module 403 is configured to determine event address information corresponding to each of the multimodal data;
[0089] The first clustering module 404 is used to cluster the multimodal data according to the event classification results and event address information to obtain at least one cluster, wherein the event classification results and event address information corresponding to various modal data in the same cluster are the same.
[0090] It should be noted that the aforementioned explanation of the embodiment of the event aggregation method based on multimodal data is also applicable to this embodiment, and will not be repeated in this embodiment.
[0091] The multimodal data-based event aggregation device of the disclosed embodiment, when processing the multimodal data to be processed, determines the event classification results and event address information corresponding to each of the multimodal data, and clusters the multimodal data to obtain at least one cluster based on the event classification results and event address information. Thus, a method for performing event aggregation on multimodal data based on the event classification results and event address information is provided, conveniently implementing event aggregation on multimodal data.
[0092] In one embodiment of the present disclosure, Figure 5 is a schematic diagram of a fifth embodiment according to the present disclosure, as shown in FIG. Figure 5 As shown, the event aggregation device 500 based on multimodal data may include: an acquisition module 501, an event classification module 502, a first determination module 503, a first clustering module 504, a second determination module 505, a mapping module 506, a generation module 507 and a second clustering module 508.
[0093] It should be noted that for a detailed description of the acquisition module 501, the first determination module 503 and the first clustering module 504, please refer to Figure 4 The description of the acquisition module 401, the first determination module 403 and the first clustering module 404 in the illustrated embodiment will not be repeated here.
[0094] In one embodiment of the present disclosure, the event classification module 502 is specifically used to: for each modal data in the multimodal data, input the modal data into the event classification model corresponding to the modal data to obtain an event classification result of the modal data.
[0095] In one embodiment of the present disclosure, multimodal data includes text modal data and visual modal data. The first determination module 503 is specifically used to: extract event addresses from the text modal data through a feature extraction model to obtain event address information corresponding to the text modal data; and determine the event address information of the visual modal data based on the position information of the camera bound to the visual modal data.
[0096] In one embodiment of the present disclosure, the apparatus 500 may further include:
[0097] The second determination module 505 is used to determine the original classification label system based on the event classification result;
[0098] The mapping module 506 is used to map the event classification results according to the mapping relationship between the original classification label system and the preset target classification label system.
[0099] In one embodiment of the present disclosure, the apparatus 500 may further include:
[0100] The generation module 507 is used to generate a cluster name for the cluster according to the event classification results and event address information corresponding to various modal data in the cluster.
[0101] In one embodiment of the present disclosure, the apparatus 500 may further include:
[0102] The second clustering module 508 is used to cluster the various modal data in the cluster cluster again according to the event subjects corresponding to the various modal data in the cluster cluster, so as to obtain at least one cluster sub-cluster corresponding to the cluster cluster, wherein the event subjects corresponding to the various modal data in the same cluster sub-cluster are the same.
[0103] It should be noted that the aforementioned explanation of the embodiment of the event aggregation method based on multimodal data is also applicable to the event aggregation device based on multimodal data in this embodiment, and will not be repeated here.
[0104] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.
[0105] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0106] like Figure 6 As shown, the electronic device 600 may include a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the device 600 may also be stored in the RAM 603. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0107] Various components in device 600 are connected to I / O interface 605, including an input unit 606, such as a keyboard, mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, optical disk, etc.; and a communication unit 609, such as a network card, modem, wireless communication transceiver, etc. The communication unit 609 allows device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0108] The computing unit 601 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 601 performs the various methods and processes described above, such as an event aggregation method based on multimodal data. For example, in some embodiments, the event aggregation method based on multimodal data can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the event aggregation method based on multimodal data described above can be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to execute the event aggregation method based on multimodal data in any other appropriate manner (for example, by means of firmware).
[0109] Various embodiments of the devices and techniques described above herein can be implemented in digital electronic circuit devices, integrated circuit devices, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), devices on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable device that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage device, at least one input device, and at least one output device, and transmit data and instructions to the storage device, the at least one input device, and the at least one output device.
[0110] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0111] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use with an instruction execution device, device or equipment or used in combination with an instruction execution device, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared or semiconductor devices, devices or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0112] To provide interaction with a user, the devices and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0113] The apparatus and techniques described herein can be implemented in a computing device that includes backend components (e.g., as a data server), or a computing device that includes middleware components (e.g., an application server), or a computing device that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the apparatus and techniques described herein), or a computing device that includes any combination of such backend components, middleware components, or front-end components. The components of the apparatus can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.
[0114] A computer device may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. The client-server relationship is established by computer programs running on the respective computers and establishing a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and VPS services ("Virtual Private Servers" or "VPS"). The server may be a cloud server, a distributed device server, or a server integrated with blockchain.
[0115] It's important to note that artificial intelligence (AI) is the study of how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). This encompasses both hardware and software technologies. AI hardware technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily encompass computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graphs.
[0116] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.
[0117] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. An event aggregation method based on multimodal data, comprising: Acquiring multimodal data to be processed, wherein the multimodal data includes textual modal data and visual modal data; For each modal data in the multimodal data, inputting the modal data into an event classification model corresponding to the modal data to obtain an event classification result for the modal data; Performing event address extraction on the text modality data through a feature extraction model to obtain event address information corresponding to the text modality data; determining event address information of the visual modality data based on location information of a camera bound to the visual modality data; Determining an original classification label system on which the event classification result is based, and mapping the event classification result according to a mapping relationship between the original classification label system and a preset target classification label system; The multimodal data is clustered according to the mapped event classification results and the event address information to obtain at least one cluster, wherein the event classification results corresponding to various modal data in the same cluster and the event address information are the same.
2. The method according to claim 1, wherein The method further comprises: For the cluster, a cluster name of the cluster is generated according to the event classification results and event address information corresponding to various modal data in the cluster.
3. The method according to claim 1, wherein The method further comprises: For the clustering cluster, the various modal data in the clustering cluster are clustered again according to the event subjects corresponding to the various modal data in the clustering cluster to obtain at least one clustering sub-cluster corresponding to the clustering cluster, wherein the event subjects corresponding to the various modal data in the same clustering sub-cluster are the same.
4. An event aggregation device based on multimodal data, comprising: An acquisition module, configured to acquire multimodal data to be processed, wherein the multimodal data includes textual modal data and visual modal data; an event classification module, configured to input, for each modal data in the multimodal data, the modal data into an event classification model corresponding to the modal data, so as to obtain an event classification result for the modal data; A first determination module is configured to extract event addresses from the text modality data using a feature extraction model to obtain event address information corresponding to the text modality data; and determine the event address information of the visual modality data based on position information of a camera bound to the visual modality data; A second determination module is used to determine the original classification label system on which the event classification result is based; A mapping module, configured to map the event classification results according to a mapping relationship between the original classification label system and a preset target classification label system; The first clustering module is used to cluster the multimodal data according to the mapped event classification results and the event address information to obtain at least one cluster, wherein the event classification results corresponding to various modal data in the same cluster and the event address information are the same.
5. The device according to claim 4, wherein The device further comprises: A generation module is used to generate a cluster name for the cluster according to event classification results and event address information corresponding to various modal data in the cluster.
6. The device according to claim 4, wherein The device further comprises: The second clustering module is used to cluster the various modal data in the cluster cluster again according to the event subjects corresponding to the various modal data in the cluster cluster, so as to obtain at least one cluster sub-cluster corresponding to the cluster cluster, wherein the event subjects corresponding to the various modal data in the same cluster sub-cluster are the same.
7. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 3.
8. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 3.
9. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 3 is implemented.
Citation Information
Patent Citations
Emergency public security event detection method based on multi-modal data
CN110232158A
Local event detection method and device, equipment and storage medium
CN113821739A