Public opinion event identification method and device, electronic equipment, storage medium and product
By obtaining image and text data in the field of food safety, and using trigger word sets and knowledge bases to enhance feature representation, the complexity of food safety public opinion data is solved, and more accurate identification and monitoring of public opinion events are achieved.
Patent Information
- Application Number
- CN202510481356.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-29
AI Technical Summary
The amount of food safety public opinion data is large and complex, and it is difficult for existing technology to effectively analyze public opinion events.
By acquiring image and text data, using the preset trigger word set and public opinion event knowledge base, the target trigger word and prompt characteristics are determined, the text and image feature representation is enhanced, and the public opinion event recognition is combined with multimodal information.
It improves the accuracy of public opinion incident identification and can better support public opinion monitoring and management in the target field.
Smart Images

Figure CN120387134A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of public opinion analysis, and particularly relates to a method, device, electronic device, storage medium and product for identifying public opinion events. Background Art
[0002] At present, food safety issues have received wide attention. With the rapid development of informatization, information spreads faster, and food safety public opinion shows an explosive growth trend. However, food safety public opinion data usually has a large volume and is more complex and diverse, and it is also difficult to effectively analyze public opinion events from these food safety public opinion data. Summary of the Invention
[0003] The present disclosure provides a method, device, electronic device, storage medium and product for identifying public opinion events.
[0004] In a first aspect, the present disclosure provides a method for identifying a public opinion event, the method comprising:
[0005] Obtaining to-be-processed public opinion data, where the to-be-processed public opinion data includes image data and text data;
[0006] Determining a target trigger word that meets semantic conditions according to the text data and a preset trigger word set, where the trigger word is used to express key information of the text semantics;
[0007] Obtaining a comprehensive text feature according to a first word feature of the target trigger word and a first text feature of the text data;
[0008] Obtaining a comprehensive image feature according to a first image feature of the image data and a target hint feature, where the target hint feature is obtained according to a public opinion event knowledge base in a target field matching the to-be-processed public opinion data, and represents image semantic guidance information in the target field;
[0009] Determining a public opinion event category corresponding to the to-be-processed public opinion data according to the comprehensive text feature and the comprehensive image feature.
[0010] In a second aspect, the present disclosure provides a device for identifying a public opinion event, the device comprising:
[0011] An obtaining module, configured to obtain to-be-processed public opinion data, where the to-be-processed public opinion data includes image data and text data;
[0012] A first determination module, configured to determine a target trigger word that meets semantic conditions according to the text data and a preset trigger word set, where the trigger word is used to express key information of the text semantics;
[0013] A first processing module, configured to obtain a comprehensive text feature according to a first word feature of the target trigger word and a first text feature of the text data;
[0014] A second processing module, configured to obtain a comprehensive image feature according to a first image feature of the image data and a target prompt feature, where the target prompt feature is obtained according to an opinion event knowledge base in a target field that matches the opinion data to be processed, and represents image semantic guidance information in the target field;
[0015] A second determination module, configured to determine an opinion event category corresponding to the opinion data to be processed according to the comprehensive text feature and the comprehensive image feature.
[0016] In a third aspect, the present disclosure provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores one or more computer programs executable by the at least one processor, and the one or more computer programs are executed by the at least one processor, so that the at least one processor can execute the above-mentioned opinion event recognition method.
[0017] In a fourth aspect, the present disclosure provides a computer-readable storage medium, on which a computer program is stored, wherein the computer program implements the above-mentioned opinion event recognition method when executed by a processor.
[0018] In a fifth aspect, the present disclosure provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above-mentioned opinion event recognition method.
[0019] For the opinion event recognition method provided by the embodiments of the present disclosure, the opinion data to be processed including image data and text data is obtained, a target trigger word that meets semantic conditions is determined according to a trigger word set and the text data, and a comprehensive text feature is obtained according to a first word feature of the target trigger word and a first text feature of the text data. In this way, the target trigger word can guide the analysis of key information in the text, enhance the text feature representation, and improve the accuracy of the text feature representation; and a comprehensive image feature is obtained according to a first image feature of the image data and a target prompt feature, which can enhance the semantic dimension for the image feature, provide target field-specific semantic guidance, and improve the accuracy of the image feature representation. Furthermore, an opinion event category corresponding to the opinion data to be processed is determined according to the comprehensive text feature and the comprehensive image feature, integrating multi-modal information of images and texts, and the enhanced image features and text features are more accurate, thereby improving the accuracy of opinion event recognition.
[0020] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood from the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The drawings are used to provide a further understanding of the present disclosure and form a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure and do not constitute a limitation to the present disclosure. By describing the detailed exemplary embodiments with reference to the drawings, the above and other features and advantages will become more apparent to those skilled in the art. In the drawings:
[0022] Figure 1 is an application scenario diagram provided for an embodiment of the present disclosure;
[0023] Figure 2 is a flowchart of a method for identifying public opinion events provided for an embodiment of the present disclosure;
[0024] Figure 3 is a schematic diagram of a method for determining target trigger words in an embodiment of the present disclosure;
[0025] Figure 4 is a principle block diagram of public opinion event identification in an embodiment of the present disclosure;
[0026] Figure 5 is a principle block diagram of obtaining target prompt features in an embodiment of the present disclosure;
[0027] Figure 6 is a block diagram of a device for identifying public opinion events provided for an embodiment of the present disclosure;
[0028] Figure 7 is a block diagram of an electronic device provided for an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] To enable those skilled in the art to better understand the technical solutions of the present disclosure, the following provides a description of the exemplary embodiments of the present disclosure with reference to the drawings. Various details of the embodiments of the present disclosure are included to assist in understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, the description herein omits the description of well-known functions and structures.
[0030] Without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other.
[0031] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0032] The terms used herein are for describing particular embodiments only and are not intended to limit the present disclosure. As used herein, the singular forms "a" and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It will also be understood that when the terms "comprises" and / or "consists of" are used in this specification, the specified features, wholes, steps, operations, elements, and / or components are present, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their groups. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect.
[0033] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in common dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the present disclosure, and will not be interpreted as having an idealized or overly formal meaning unless expressly so defined herein.
[0034] In the technical solution of the present disclosure, the processing of the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information complies with the provisions of relevant laws and regulations and does not violate public order and good customs. The use of the user data in this technical solution follows the relevant national laws and regulations (for example, "Information Security Technology - Personal Information Security Specification", etc.). For example, corresponding regulatory measures are taken for personal information access control; regulations are imposed on the display of personal information; the purpose of using personal information does not exceed the direct or reasonable association scope; when using personal information, the clear identity indication is eliminated to avoid precise positioning to a specific individual.
[0035] The public opinion analysis in the field of food safety or other fields is very important. The accurate identification of public opinion events can provide a basis for relevant departments or enterprises to timely discover and properly handle crises, and also helps the regulatory authorities to formulate more targeted and forward-looking regulatory strategies, improve regulatory efficiency, and thus protect the relevant rights and interests of the public. However, public opinion data usually has a large volume and is more complex and diverse, which makes it more difficult to accurately and efficiently extract valuable information from it and identify public opinion events.
[0036] In the embodiment of the present disclosure, for the method for identifying public opinion events, for the to-be-processed public opinion data, the target trigger words that meet the semantic conditions can be determined according to the preset trigger word set and text data. According to the first word feature of the target trigger words and the first text feature of the text data, the comprehensive text feature is obtained. In this way, through the target trigger words, the text feature can be enhanced, the key information of the text can be guided and analyzed, and the accuracy of the text feature representation can be improved. And based on the target prompt feature of the target field and the first image feature of the image data, the comprehensive image feature can be obtained. In this way, the semantic dimension can be enhanced for the image feature, the semantic guidance targeted at the target field can be provided, and the accuracy of the image feature representation can be improved. Furthermore, according to the comprehensive text feature and the comprehensive image feature, the public opinion event category corresponding to the to-be-processed public opinion data is determined. In this way, the multi-modal information of the image and the text can be integrated, and the image feature and the text feature are more accurate, thereby improving the accuracy of the public opinion event identification and providing stronger support for the public opinion monitoring in the target field.
[0037] Of course, it should be noted that the target field in the embodiment of the present disclosure is not limited to the food safety field, and it is also applicable to the public opinion data analysis in other fields. And in the embodiment of the present disclosure, it is not limited to the identification of public opinion events, and it is also applicable to the identification of other events. The embodiment of the present disclosure does not limit this.
[0038] Figure 1 Schematically shows the application scenario diagram of the method and device for identifying public opinion events provided by the embodiment of the present disclosure.
[0039] As Figure 1 shown, the application scenario of the embodiment of the present disclosure may include a terminal device 101, a network 103, and a server 102. The network 103 is used to provide a medium for a communication link between the terminal device 101 and the server 102. The network 103 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0040] Users can use the terminal device 101 to interact with the server 102 through the network 103 to receive or send messages, etc. Various communication client applications may be installed on the terminal device 101, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0041] The terminal device 101 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0042] Server 102 may be a server that provides various services. For example, it may be a background management server (merely an example) that supports websites browsed by users using terminal device 101. The background management server may analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device. For example, a user may input the public opinion data to be processed through terminal device 101. Terminal device 101 sends a public opinion event recognition request including the public opinion data to be processed to server 102. Server 102 determines the recognized public opinion event category based on the method in the embodiments of the present disclosure and returns it to terminal device 101. Terminal device 101 may display the obtained public opinion event category to the user through an interface.
[0043] It should be noted that the public opinion event recognition method and device provided in the embodiments of the present disclosure may be executed by server 102. Correspondingly, the public opinion event recognition method and device provided in the embodiments of the present disclosure may be set in server 102. The public opinion event recognition method and device provided in the embodiments of the present disclosure may also be executed by a server or a server cluster different from server 102 and capable of communicating with terminal device 101 and / or server 102. Correspondingly, the public opinion event recognition method and device provided in the embodiments of the present disclosure may also be set in a server or a server cluster different from server 102 and capable of communicating with terminal device 101 and / or server 102.
[0044] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in
[0045] Figure 2 is a flowchart of a public opinion event recognition method provided in an embodiment of the present disclosure. Referring to Figure 2 , the method includes:
[0046] S210: Obtain the public opinion data to be processed, where the public opinion data to be processed includes image data and text data.
[0047] In the embodiments of the present disclosure, it can be applied to public opinion data analysis in any field. For example, when it is necessary to conduct public opinion analysis or monitoring in a certain field, the public opinion data to be processed can be obtained through web crawling by keywords or the like. The present disclosure does not limit the acquisition method of the public opinion data to be processed.
[0048] And usually some information and other include graphic and text types. In the embodiments of the present disclosure, when conducting public opinion event analysis, the public opinion data to be processed including graphics and text can be obtained. The combination of graphics and text can usually describe more information.
[0049] S220: Determine a target trigger word that meets the semantic conditions according to the text data and a preset trigger word set, where the trigger word is used to express key information of the text semantics.
[0050] In the embodiments of the present disclosure, the trigger word can be understood as usually being a keyword in the text, some words that can express the key information of the text.
[0051] S230: Obtain a comprehensive text feature according to the first word feature of the target trigger word and the first text feature of the text data.
[0052] S240: Obtain a comprehensive image feature according to the first image feature of the image data and the target prompt feature, where the target prompt feature is obtained according to the knowledge base of public opinion events in the target field that matches the to-be-processed public opinion data, and represents the image semantic guidance information of the target field.
[0053] S250: Determine the category of the public opinion event corresponding to the to-be-processed public opinion data according to the comprehensive text feature and the comprehensive image feature.
[0054] In the embodiments of the present disclosure, for the to-be-processed public opinion data, the target trigger word and the target prompt feature can be combined to enhance the text feature representation and the image feature representation respectively, improve the accuracy of the text feature and the image feature, better meet the feature requirements of the target field that matches the to-be-processed public opinion data, reduce semantic understanding deviation, and then determine the category of the public opinion event corresponding to the to-be-processed public opinion data according to the enhanced comprehensive text feature and comprehensive image feature, improving the accuracy of public opinion event recognition.
[0055] The following elaborates on the public opinion event recognition method according to the embodiments of the present disclosure.
[0056] As mentioned above, for the above-mentioned step S220, the present disclosure also provides a specific implementation manner.
[0057] Refer to Figure 3 As shown, it is a schematic diagram of the method for determining the target trigger word in the embodiments of the present disclosure. For ease of explanation, the above-mentioned step S220 will be introduced in combination with Figure 3 to introduce the above-mentioned step S220.
[0058] In a possible embodiment, determining a target trigger word that meets the semantic conditions according to the text data and a preset trigger word set includes:
[0059] S1: Generate a first word vector of the trigger words in the trigger word set based on the preset trigger word set.
[0060] For example, the trigger word set can be an existing data set, which is not limited thereto, such as Figure 3As shown, the trigger words in the trigger word set can be input into the language model, and based on the language model, the first word vector of each trigger word can be generated. Here, the language model can be, for example, the encoder representation of the Bidirectional Encoder Representations from Transformers (BERT), but the embodiments of the present disclosure do not limit this.
[0061] For example, the trigger word in the preset trigger word set is w i , and based on the BERT model, the first word vector h i of each trigger word w i is generated. Then, the first word vectors corresponding to the multiple trigger words output by the BETR model can be represented as H = [h1, h2, … h n , where h i ∈ R d , and d is the dimension of the first word vector.
[0062] S2: Perform part-of-speech tagging on the text data, extract the words labeled with the target part of speech from the text data, and generate the second word vector of the extracted words.
[0063] Among them, the target part of speech can be a verb or other types of parts of speech. The embodiments of the present disclosure do not limit this and can be set according to experience and requirements.
[0064] Generally, parts of speech include verbs, nouns, adjectives, etc. In one possible embodiment, part-of-speech tagging can be performed based on preset rules. For example, the part-of-speech rules can be set based on linguistic knowledge. In another possible embodiment, statistical or deep learning methods can be used to train a model based on a corpus with labeled parts of speech for part-of-speech tagging, etc. The specific part-of-speech tagging method is not limited in the embodiments of the present disclosure.
[0065] Refer to Figure 3 As shown, perform part-of-speech tagging on the text data and extract the words with the target part of speech. Among them, the extracted words may be one or more. For example, the set of extracted words is S new = [v1, v2, …, v n , where v i is the i-th word in the set of extracted words. Input the extracted words into the language model, such as the BERT model. The BERT model outputs the corresponding second word vector. For example, the second word vector corresponding to each v i is H i . Then, the output second word vector can be represented as H new = [H1, H2, …, H n , where H i ∈ R d。
[0066] S3: Determine the similarity between the extracted word and the trigger words in the trigger word set according to the first word vector and the second word vector.
[0067] In the embodiments of the present disclosure, the similarity can be determined by calculating the cosine similarity and / or Euclidean distance between vectors, etc., and there is no limitation thereto.
[0068] In a possible embodiment, determine the cosine similarity between the first word vector and the second word vector, and determine the Euclidean distance between the first word vector and the second word vector. According to the cosine similarity and the Euclidean distance, determine the similarity between the extracted word and the trigger words in the trigger word set.
[0069] For example, calculate the second word vector H i and the first word vector h j The cosine similarity between them is CoS(H i , h j ), calculate the Euclidean distance as EuD(H i , h j ), and perform weighted synthesis to obtain the final similarity as WeightSim(H i , h j ), then it can be specifically expressed as:
[0070]
[0071] WeightSim(H i , h j ) = w1·CoS(H i , h j ) + w2·NormEuD(H i , h j )
[0072]
[0073] Among them, w1 and w2 are weight coefficients, and there is no limitation in the embodiments of the present disclosure.
[0074] In this way, by performing weighted synthesis on the similarities calculated in different ways to obtain the final similarity, the accuracy of the final similarity can be improved.
[0075] S4: Determine the candidate trigger words according to the similarity and the word knowledge base in the target domain.
[0076] For this step, in a possible embodiment, 1) According to the similarity, screen out the initial trigger words that meet the similarity conditions from the extracted words and the trigger word set.
[0077] For example, the similarity between the extracted word v i and the trigger word w in the trigger word set j is passed through a linear layer, where the weight matrix of the linear layer is w ∈ R m×k , and the bias term is b ∈ R k , k is the output dimension of the linear layer, and the transformed vector is z ij . A threshold θ can be set to determine whether the similarity condition is met. For example, the judgment function is f(z ij ), then specifically:
[0078]
[0079] Among them, if f(z ij ) = 1, it means that the similarity between the i-th word in the extracted word set and the j-th trigger word in the trigger word set is greater than the threshold, that is, the similarity condition is met, and it can be determined as the initial trigger word. For example, the finally obtained initial trigger word set can be expressed as: T = {t ij |f(z ij ) = 1, i = 1, 2,..., n, j = 1, 2...k}.
[0080] 2) According to the word knowledge base in the target domain, obtain the extended words that meet the association conditions with the initial trigger words.
[0081] In the embodiments of the present disclosure, according to the knowledge base in the target domain, the part of speech, semantic role, and context information of the initial trigger words can be analyzed, and the extended words that are relevant and semantically similar to the initial trigger words can be screened out.
[0082] For example, as Figure 3 shown, the initial trigger word is T = {t ij}, and the extended words that meet the association conditions with each initial trigger word are obtained from the knowledge base. There can be multiple extended words, and then the extended words can be expressed as {E1, E2,...E n}, where represents the set of extended words that meet the association conditions with the initial trigger word t i , and m i is the number of corresponding extended words.
[0083] 3) According to the initial trigger words and the extended words, obtain the candidate trigger words.
[0084] For example, the i-th extended word E i and the initial trigger word t i are merged to obtain the candidate trigger word C i , that is, C i = t i ∪ E i .
[0085] S5: Based on a preset prompt statement template, perform semantic analysis on the candidate trigger words, and determine target trigger words that meet the semantic conditions from the candidate trigger words.
[0086] In the embodiments of the present disclosure, the prompt statement template represents a statement sample including at least one preset filling position, which can be set in advance according to requirements, and the prompt statement template can be set to one or more, and the embodiments of the present disclosure do not limit this.
[0087] For example, in the field of food safety, the food safety domain knowledge base covers various aspects of information such as the characteristics of various foods, production and processing processes, quality standards, and common safety problems, etc., and a large and orderly semantic association network can be constructed. For example, the knowledge base records professional knowledge such as the scope of use and limit standards of different food additives, the impact of various pathogenic microorganisms on food safety, etc. In this way, in the embodiments of the present disclosure, the prompt statement template can be set according to the food safety domain knowledge base and requirements. The prompt statement template can guide the model to more easily capture key semantic information related to food safety events. For example, in a text related to a food recall event, the preset filling positions in the prompt statement template can be the reason position, the food batch position, and the recall scope position. In this way, the prompt statement template can guide the model to focus on elements such as the recall reason, the food batches involved, and the recall scope.
[0088] In the embodiments of the present disclosure, possible implementation manners are provided. 1) Based on a preset prompt statement template, fill the candidate trigger words into the corresponding filling positions in the prompt statement template to obtain a prompt statement. The prompt statement template includes at least one preset filling position.
[0089] For example, the prompt statement template is: A certain product fails to pass the safety inspection because of _ (reason filling position), and then fill the candidate trigger word into this filling position to obtain a complete prompt statement.
[0090] 2) Based on the model, use the prompt statement as the input, and perform semantic analysis on the prompt statement to obtain the first semantic score of the prompt statement.
[0091] In the embodiments of the present disclosure, the model can be pre-trained. This model is used for semantic analysis. For example, it can analyze the matching degree between the candidate trigger words and the prompt statement template, and can also analyze whether the statement is complete, correct, and analyze its coherence and logic, etc., so as to obtain the first semantic score of the prompt statement.
[0092] 3) Determine the candidate trigger words corresponding to the prompt statements whose first semantic scores meet the score conditions as the target trigger words.
[0093] For example, the score condition may be that the first semantic score is greater than or equal to a preset threshold, which is not limited in the embodiments of the present disclosure.
[0094] In this way, in the embodiments of the present disclosure, the word knowledge base of the target domain can be used as the basis for expanding the initial trigger words, and the model can be guided to understand and analyze the text from a set perspective based on the prompt statement template, which can enhance the semantic understanding of the target domain, and the obtained target trigger words are more in line with the requirements of the target domain, thereby improving the accuracy of subsequent identification of public opinion events in the target domain.
[0095] Refer to Figure 4 As shown, it is a principle block diagram of public opinion event recognition in the embodiments of the present disclosure. For ease of explanation, the above steps S230 - S250 will be introduced in conjunction with Figure 4 the above.
[0096] In a possible embodiment, for the above step S230, the embodiments of the present disclosure include:
[0097] 1) Extract features from the text data to obtain the first text feature of the text data.
[0098] For example, refer to Figure 4 As shown, the first text feature V can be obtained by extracting features through a Transformer encoding model T ={v t1 ,v t2 ,…,v tm}, where m is the length of the text sequence, and d t is the dimension of the feature vector of the first text feature.
[0099] 2) Extract features from the target trigger word to obtain the first word feature of the target trigger word, and the feature vector dimensions of the first word feature and the first text feature are the same.
[0100] For example, refer to Figure 4 As shown, the first word feature V is obtained by extracting features of the target trigger word through a Transformer encoding model W ={v w1 ,v w2 ,…,v wn}, n is the length of the target trigger word sequence, and the feature vector dimensions of the first word feature and the first text feature are the same.
[0101] 3) Concatenate the first text feature and the first word feature to obtain a comprehensive text feature.
[0102] For example, through the concatenation operation, the obtained comprehensive text feature can be expressed as:
[0103] V new = {v t1 , v t2 , …, v tm , v w1 , v w2 , …, v wn}.
[0104] In this way, the target trigger word and the features of the text data are fused, and the target trigger word can express the key information of the text semantics in the target field, so that the fused comprehensive text features can better represent the key semantic information hidden in the text, enhancing the text feature representation and improving the accuracy of the text feature representation.
[0105] For the above step S240, the present disclosure provides possible implementation manners. According to the first image feature and the target prompt feature of the image data, a comprehensive image feature is obtained, including:
[0106] 1) The image data is segmented into multiple image patches, and feature extraction is performed on each image patch to obtain a sub-graph vector of each image patch.
[0107] 2) According to the sub-graph vectors of each image patch, the first image feature of the image data is obtained.
[0108] In the embodiments of the present disclosure, in order to improve the efficiency and accuracy of image feature extraction, the image data can be segmented into multiple image patches and then feature extraction is performed on each image patch.
[0109] For example, if the image data is I, it is segmented into N image patches {p1, p2, …, p N} and an embedding operation is performed on each image patch p i to obtain a sub-graph vector e i , which can be expressed as e i = PatchEmbed(p i ), where PatchEmbed is an image patch embedding function, e i ∈ R D , D is the vector dimension. After being processed by the self-attention module and the feed-forward neural network of L layers, assuming the output of the L-th layer is Then: Then the finally obtained first image feature of the image data is Where D’ is the dimension of the final feature vector.
[0110] Among them, the acquisition of the first image feature can be achieved through an image processing model, such as a vision transformer (VIT) model. By inputting image data into the VIT model, the first image feature output by the VIT model can be obtained. Specifically, the image processing model is not limited in the embodiments of the present disclosure.
[0111] 3) Obtaining a target prompt feature, and performing a splicing operation on the target prompt feature and the first image feature to obtain a comprehensive image feature.
[0112] For example, Figure 4 As shown, the target prompt feature and the first image feature can be spliced together, that is, feature fusion is achieved to obtain a comprehensive image feature.
[0113] In the disclosed embodiment, a pre-prompt strategy based on the target prompt feature is introduced to fuse the target prompt feature with the first image feature, which can add a semantic dimension to the image feature representation and improve the expressive ability in semantic understanding. The target prompt feature can be understood as a kind of image semantic guidance information to assist image features in learning more discriminative feature representations. For example, food spoilage events may appear in images as color changes and abnormal textures. The pre-prompt strategy can provide targeted semantic guidance information for image features based on prior knowledge of event categories, such as features learned from historical food spoilage events, thereby improving the accuracy of image feature representation.
[0114] With respect to the above step S250, the present disclosure provides a possible implementation method for determining the public opinion event category corresponding to the public opinion data to be processed based on the comprehensive text features and comprehensive image features, including:
[0115] 1) Perform a weighted sum operation based on the comprehensive text features and the comprehensive image features to obtain the first feature.
[0116] Furthermore, in the embodiment of the present disclosure, the vector dimensions of the integrated text features and the integrated image features may be different, so it is also necessary to perform unified processing of the vector dimensions, for example, by using an MLP module, the parameters of the MLP module are θ' MLP , then the processed comprehensive text features can be expressed as: The processed comprehensive image features can be expressed as:
[0117] Furthermore, Figure 4 As shown, the processed comprehensive text features and comprehensive image features are input into the dynamic gate module, wherein the dynamic gate module may include an image dynamic gate G img and text dynamic gate G text, which are respectively used to control the weights of image features and text features in the fusion. For example, the feature vector obtained through the dynamic gate module is V fusion , and then through a Multilayer Perceptron (MLP). For example, the function of the MLP is MLP θ' , the first feature obtained through the MLP can be expressed as:
[0118] In the embodiments of the present disclosure, the weighted parameters of the image features and text features in the dynamic gate module can be determined through pre-training. For example, during training, according to the error between the predicted event category and the true event category obtained from the training fusion features after the corresponding fusion of the image and text, the parameters in the dynamic gate module can be adjusted using the backpropagation algorithm to optimize the fusion effect. In each round of training, the dynamic gate module can dynamically adjust the attention mechanism and fusion strategy according to the current image features, text features, and the learning state of the model, so as to fully exploit the complementarity between multi-modal information, improve the quality and accuracy of the fused feature vector. Thus, the dynamic gate module can, based on its internal fusion mechanism and parameter adaptive adjustment strategy, fuse and optimize the input image features and text features, thereby generating a multi-modal information feature fused by weighted summation and a feature vector with a unified dimension.
[0119] 2) Perform a non-linear mapping on the first feature to obtain the probability value belonging to each preset category.
[0120] For example, as Figure 4 shown, the first feature can be passed through an activation function, and the activation function performs a non-linear mapping on the first feature. For example, the first feature is V output , and through the activation function mapping, V sigmoid =σ(V output ), and the value of each element after mapping is between (0, 1). The value of each element represents the probability value belonging to each preset category.
[0121] Among them, the preset categories of public opinion events can be set in advance. Through the learning and training of the activation function in advance, the activation function can be used to identify the probability value belonging to each preset category. For example, in the field of food safety, the preset categories of food safety public opinion events can be set to include food unqualified, food unhygienic, etc. The embodiments of the present disclosure do not limit this.
[0122] 3) According to the probability value belonging to each preset category and the preset judgment rule, determine the public opinion event category corresponding to the to-be-processed public opinion data.
[0123] For example, the preset judgment rule is R. According to V sigmoidGiven the judgment rule R, if the corresponding public opinion event category is y, it can be expressed as y = Classify(V sigmoid , R), where the preset judgment rule can be the category corresponding to the maximum probability value as the recognized public opinion event category.
[0124] In this way, in the embodiments of the present disclosure, the text features are enhanced based on the target trigger word to obtain comprehensive text features, and the image features are enhanced based on the target prompt features to obtain comprehensive image features. Then, based on the comprehensive text features and comprehensive image features, fusion is performed to determine the public opinion event category corresponding to the to-be-processed public opinion data. This not only enables the recognition of public opinion events by fusing multi-modal information of text and images, but also makes the text features and image features more accurate, thereby improving the recognition accuracy of public opinion event categories in the target field.
[0125] In the embodiments of the present disclosure, a pre-prompt strategy is used to enhance the image features. Among them, for the method of obtaining the target prompt features in the pre-prompt strategy, corresponding possible embodiments are also provided. Refer to Figure 5 As shown, it is a principle block diagram for obtaining the target prompt features in the embodiments of the present disclosure. Specifically, obtaining the target prompt features includes:
[0126] 1) Obtain the training image data and training text data of public opinion events in the knowledge base of public opinion events in the target field.
[0127] 2) Extract features from the training image data to obtain second image features, and extract features from the training text data to obtain second text features.
[0128] For example, as Figure 5 shown, taking the food safety field as an example of the target field, the training image data is I food , which is input into the VIT model to extract features from the training image data to obtain second image features which can be expressed as: where N represents the number of image patches into which the training image data is segmented, is the vector dimension of the second image feature.
[0129] The training text data is T food , and based on the Transformer model, features are extracted from the training text data to generate second text features which can be expressed as: where N represents the length of the training text data, is the vector dimension of the second text feature.
[0130] 3) Concatenate the second image feature and the first hint feature of the current iteration to obtain a comprehensive hint feature, where the initial hint feature of the first iteration is randomly generated.
[0131] In the embodiments of the present disclosure, the target hint feature can be obtained through multiple rounds of iterative learning. For example, the initial hint feature of the first iteration can be randomly selected from a normal distribution N(μ,σ 2 ), and the initial hint feature can be represented as V prompt ~N(μ,σ 2 ), where the present disclosure does not limit this.
[0132] Taking the first iteration as an example, concatenate the initial hint feature and the second image feature to obtain a comprehensive hint feature, which can be represented as It can be seen that the vector dimension after the concatenation operation changes to V fused ,
[0133]
[0134] 4) Determine the first loss value according to the comprehensive hint feature and the second text feature.
[0135] In the embodiments of the present disclosure, when determining the first loss value according to the comprehensive hint feature and the second text feature, the vector dimensions of the comprehensive hint feature and the second text feature may be different. Therefore, it is also necessary to perform a unified processing of the vector dimensions to transform the comprehensive hint feature V fused into a feature with the same dimension as the second text feature For example, an MLP module can be used. The parameters of the MLP module are θ
[0136] , and a vector with the same dimension as the second text feature is generated based on the comprehensive hint feature MLP It can be represented as where By making it have the same vector dimension as the text feature, it is convenient for subsequent loss calculation and optimization adjustment. Then, according to the second text feature and the transformed comprehensive hint feature the first loss value is determined.
[0137] 5) Adjust the first hint feature according to the first loss value to obtain the second hint feature of the next iteration until the iteration end condition is met, and obtain the target hint feature of the final iteration.
[0138] In the embodiments of the present disclosure, through continuous iterative learning, the first loss value can be calculated based on the backpropagation method, and the parameters can be updated to continuously optimize and adjust the prompt features. After each round of iteration, the association weights of the words in the prior knowledge graph can also be adjusted according to the recognition results of the public opinion events corresponding to the current public opinion data and the feedback of the model on the use of the prompt features. If a certain word appears frequently in the successfully recognized public opinion events and plays a key role in the recognition results, its weight in the prompt vector is increased; conversely, if a certain word causes recognition errors or has a low correlation with the actual public opinion events, its weight is decreased. At the same time, an adaptive learning rate strategy is introduced to dynamically adjust the learning rate according to the optimization effect of the prompt vector to ensure the stability and effectiveness of the optimization process, so that the final target prompt vector is more generalized in the field of food safety public opinion. Furthermore, when recognizing public opinion events, integrating the pre-trained template prompt features can improve the accuracy of recognizing public opinion events in the target field.
[0139] It can be understood that the above-mentioned method embodiments mentioned in the present disclosure can be combined with each other to form combined embodiments without violating the principle logic. Due to space limitations, the present disclosure will not elaborate further. Those skilled in the art can understand that in the above methods of the specific implementation manner, the specific execution order of each step should be determined according to its function and possible internal logic.
[0140] In addition, the present disclosure also provides a public opinion event recognition device, an electronic device, a computer-readable storage medium, and a computer program product, all of which can be used to implement any public opinion event recognition method provided by the present disclosure. The corresponding technical solutions and descriptions are referred to the corresponding records in the method part and will not be elaborated further.
[0141] Figure 6 It is a block diagram of a public opinion event recognition device provided by an embodiment of the present disclosure.
[0142] Referring to Figure 6 , an embodiment of the present disclosure provides a public opinion event recognition device, which includes:
[0143] An acquisition module 61, configured to acquire the to-be-processed public opinion data, where the to-be-processed public opinion data includes image data and text data;
[0144] A first determination module 62, configured to determine a target trigger word that meets the semantic conditions according to the text data and a preset trigger word set, where the trigger word is used to express the key information of the text semantics;
[0145] A first processing module 63, configured to obtain a comprehensive text feature according to the first word feature of the target trigger word and the first text feature of the text data;
[0146] A second processing module 64, configured to obtain a comprehensive image feature according to a first image feature of the image data and a target prompt feature, where the target prompt feature is obtained according to a knowledge base of public opinion events in a target field that matches the to-be-processed public opinion data, and represents image semantic guidance information of the target field;
[0147] A second determination module 65, configured to determine a public opinion event category corresponding to the to-be-processed public opinion data according to the comprehensive text feature and the comprehensive image feature.
[0148] In an optional embodiment, when determining a target trigger word that meets semantic conditions according to the text data and a preset trigger word set, the first determination module 62 is configured to:
[0149] Generate a first word vector of trigger words in the trigger word set based on the preset trigger word set;
[0150] Perform part-of-speech tagging on the text data, extract words labeled with a target part of speech from the text data, and generate a second word vector of the extracted words;
[0151] Determine the similarity between the extracted words and the trigger words in the trigger word set according to the first word vector and the second word vector;
[0152] Determine candidate trigger words according to the similarity and a knowledge base of words in the target field;
[0153] Perform semantic analysis on the candidate trigger words based on a preset prompt statement template, and determine a target trigger word that meets semantic conditions from the candidate trigger words.
[0154] In an optional embodiment, when determining candidate trigger words according to the similarity and a knowledge base of words in the target field, the first determination module 62 is configured to:
[0155] Screen out initial trigger words that meet the similarity condition from the extracted words and the trigger word set according to the similarity;
[0156] Obtain extended words that meet the association condition with the initial trigger words according to the knowledge base of words in the target field;
[0157] Obtain candidate trigger words according to the initial trigger words and the extended words.
[0158] In an optional embodiment, when performing semantic analysis on the candidate trigger words based on a preset prompt statement template and determining a target trigger word that meets semantic conditions from the candidate trigger words, the first determination module 62 is configured to:
[0159] Based on a preset prompt statement template, fill the candidate trigger word into the corresponding filling position in the prompt statement template to obtain a prompt statement, where the prompt statement template includes at least one preset filling position;
[0160] Based on a model, use the prompt statement as input, and perform semantic analysis on the prompt statement to obtain a first semantic score of the prompt statement;
[0161] Determine the candidate trigger word corresponding to the prompt statement whose first semantic score meets the score condition as the target trigger word.
[0162] In an alternative embodiment, when obtaining the comprehensive text feature according to the first word feature of the target trigger word and the first text feature of the text data, the first processing module 63 is configured to:
[0163] Extract features from the text data to obtain a first text feature of the text data;
[0164] Extract features from the target trigger word to obtain a first word feature of the target trigger word, where the feature vector dimensions of the first word feature and the first text feature are the same;
[0165] Perform a splicing operation on the first text feature and the first word feature to obtain a comprehensive text feature.
[0166] In an alternative embodiment, when obtaining the comprehensive image feature according to the first image feature of the image data and the target prompt feature, the second processing module 64 is configured to:
[0167] Divide the image data into multiple image blocks, and extract features from each image block to obtain a sub-image vector of each image block;
[0168] Obtain a first image feature of the image data according to the sub-image vector of each image block;
[0169] Obtain a target prompt feature, and perform a splicing operation on the target prompt feature and the first image feature to obtain a comprehensive image feature.
[0170] In an alternative embodiment, when obtaining the target prompt feature, the second processing module 64 is configured to:
[0171] Obtain the training image data and training text data of the public opinion events in the knowledge base of public opinion events in the target field;
[0172] Extract features from the training image data to obtain a second image feature, and extract features from the training text data to obtain a second text feature;
[0173] Performing a splicing operation on the second image feature and the first prompt feature of the current iteration to obtain a comprehensive prompt feature, wherein the initial prompt feature of the first iteration is randomly generated;
[0174] determining a first loss value according to the comprehensive prompt feature and the second text feature;
[0175] According to the first loss value, the first prompt feature is adjusted to obtain the second prompt feature of the next iteration, until the iteration end condition is met, and the target prompt feature of the final iteration is obtained.
[0176] In an optional embodiment, when determining the public opinion event category corresponding to the public opinion data to be processed based on the comprehensive text features and the comprehensive image features, the second determination module 65 is configured to:
[0177] Performing a weighted sum operation based on the comprehensive text feature and the comprehensive image feature to obtain a first feature;
[0178] Performing nonlinear mapping on the first feature to obtain a probability value belonging to each preset category;
[0179] According to the probability value belonging to each preset category and the preset judgment rules, the public opinion event category corresponding to the public opinion data to be processed is determined.
[0180] Each module in the above-mentioned public opinion event identification device can be implemented in whole or in part through software, hardware, or a combination thereof. Each of the above modules can be embedded in or independent of the processor of the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each of the above modules.
[0181] Figure 7 A block diagram of an electronic device provided in an embodiment of the present disclosure.
[0182] Reference Figure 7 An embodiment of the present disclosure provides an electronic device, which includes: at least one processor 701; at least one memory 702, and one or more I / O interfaces 703, connected between the processor 701 and the memory 702; wherein the memory 702 stores one or more computer programs that can be executed by the at least one processor 701, and the one or more computer programs are executed by the at least one processor 701 so that the at least one processor 701 can execute the above-mentioned public opinion event identification method.
[0183] Each module in the above-mentioned electronic device can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.
[0184] An embodiment of the present disclosure also provides a computer-readable storage medium, on which a computer program is stored, wherein the computer program, when executed by a processor, implements the above-mentioned public opinion event recognition method. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.
[0185] An embodiment of the present disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in the processor of an electronic device, the processor in the electronic device executes the above-mentioned public opinion event recognition method.
[0186] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations. In the hardware implementation, the division of the functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component can have multiple functions, or a function or step can be executed by several physical components in cooperation. Some or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or be implemented as hardware, or be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable storage medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium).
[0187] As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information such as computer-readable program instructions, data structures, program modules, or other data. Computer storage media includes but is not limited to random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technologies, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. Additionally, as is well known to those of ordinary skill in the art, communication media typically contains computer-readable program instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.
[0188] The computer-readable program instructions described herein can be downloaded to each computing / processing device from a computer-readable storage medium or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0189] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.
[0190] The computer program product described herein may be implemented specifically in the form of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is specifically embodied as a computer storage medium. In another alternative embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.
[0191] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer - readable program instructions.
[0192] These computer-readable program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, create an apparatus for implementing the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture including instructions that implement various aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0193] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0194] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, and the module, segment of code, or portion of an instruction may include one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by special-purpose hardware-based systems that perform the specified functions or acts, or by combinations of special-purpose hardware and computer instructions.
[0195] Example embodiments have been disclosed herein, and although specific terms are employed, they are used in a generic and descriptive sense only and not for purposes of limitation. In some instances, it will be apparent to those skilled in the art that, unless otherwise expressly stated, the features, characteristics, and / or elements described in connection with a particular embodiment may be used singly or in combination with other embodiments. Accordingly, those skilled in the art will appreciate that various forms and details may be changed without departing from the scope of the present disclosure as set forth by the appended claims.
Claims
1. A method for identifying public opinion events, characterized in that, Including: Obtain the public opinion data to be processed, where the public opinion data to be processed includes image data and text data; According to the text data and a preset trigger word set, determine a target trigger word that meets the semantic conditions, where the trigger word is used to express the key information of the text semantics; According to the first word feature of the target trigger word and the first text feature of the text data, obtain a comprehensive text feature; According to the first image feature of the image data and the target prompt feature, obtain a comprehensive image feature, where the target prompt feature is obtained according to the public opinion event knowledge base in the target field that matches the public opinion data to be processed, and represents the image semantic guidance information in the target field; According to the comprehensive text feature and the comprehensive image feature, determine the category of the public opinion event corresponding to the public opinion data to be processed.
2. The method according to claim 1, characterized in that, The step of determining a target trigger word that meets the semantic conditions according to the text data and a preset trigger word set includes: Based on a preset trigger word set, generate the first word vector of the trigger words in the trigger word set; Perform part-of-speech tagging on the text data, and extract the words labeled with the target part of speech from the text data to generate the second word vector of the extracted words; According to the first word vector and the second word vector, determine the similarity between the extracted words and the trigger words in the trigger word set; According to the similarity and the word knowledge base in the target field, determine candidate trigger words; Based on a preset prompt statement template, perform semantic analysis on the candidate trigger words, and determine the target trigger words that meet the semantic conditions from the candidate trigger words.
3. The method according to claim 2, wherein The step of determining candidate trigger words according to the similarity and the word knowledge base in the target field includes: According to the similarity, screen out the initial trigger words that meet the similarity conditions from the extracted words and the trigger word set; According to the word knowledge base in the target field, obtain the extended words that meet the association conditions with the initial trigger words; According to the initial trigger words and the extended words, obtain candidate trigger words.
4. The method according to claim 2, wherein The step of performing semantic analysis on the candidate trigger words based on a preset prompt statement template and determining the target trigger words that meet the semantic conditions from the candidate trigger words includes: Based on a preset prompt statement template, fill the candidate trigger words into the corresponding filling positions in the prompt statement template to obtain a prompt statement, where the prompt statement template includes at least one preset filling position; Based on a model, use the prompt statement as the input, perform semantic analysis on the prompt statement, and obtain the first semantic score of the prompt statement; Determine the candidate trigger words corresponding to the prompt statements whose first semantic scores meet the score conditions as the target trigger words.
5. The method according to any one of claims 1 to 4, characterized in that The step of obtaining a comprehensive text feature according to the first word feature of the target trigger word and the first text feature of the text data includes: Perform feature extraction on the text data to obtain the first text feature of the text data; Perform feature extraction on the target trigger word to obtain the first word feature of the target trigger word, where the feature vector dimensions of the first word feature and the first text feature are the same; Perform a splicing operation on the first text feature and the first word feature to obtain a comprehensive text feature.
6. The method according to any one of claims 1-4, characterized in that, Obtaining a comprehensive image feature according to the first image feature of the image data and the target prompt feature includes: Segment the image data into multiple image patches, and perform feature extraction on each image patch to obtain a sub-image vector for each image patch; Obtain the first image feature of the image data according to the sub-image vector of each image patch; Obtain a target prompt feature, and perform a splicing operation on the target prompt feature and the first image feature to obtain a comprehensive image feature.
7. The method according to claim 6, characterized in that, The obtaining of the target prompt feature includes: Obtain the training image data and training text data of the public opinion events in the public opinion event knowledge base of the target domain; Perform feature extraction on the training image data to obtain a second image feature, and perform feature extraction on the training text data to obtain a second text feature; Perform a splicing operation on the second image feature and the first prompt feature of the current round of iteration to obtain a comprehensive prompt feature, where the initial prompt feature of the first round of iteration is randomly generated; Determine a first loss value according to the comprehensive prompt feature and the second text feature; Adjust the first prompt feature according to the first loss value to obtain the second prompt feature of the next round of iteration until the iteration end condition is met, and obtain the target prompt feature of the final round of iteration.
8. The method according to claim 1, characterized in that, Determining the public opinion event category corresponding to the to-be-processed public opinion data according to the comprehensive text feature and the comprehensive image feature includes: Perform a weighted summation operation on the comprehensive text feature and the comprehensive image feature to obtain a first feature; Perform a non-linear mapping on the first feature to obtain probability values belonging to each preset category; Determine the public opinion event category corresponding to the to-be-processed public opinion data according to the probability values belonging to each preset category and a preset judgment rule.
9. An apparatus for identifying public opinion events, characterized in that, including: An acquisition module for acquiring to-be-processed public opinion data, where the to-be-processed public opinion data includes image data and text data; A first determination module for determining a target trigger word that meets the semantic condition according to the text data and a preset trigger word set, where the trigger word is used to express key information of the text semantics; A first processing module for obtaining a comprehensive text feature according to the first word feature of the target trigger word and the first text feature of the text data; A second processing module for obtaining a comprehensive image feature according to the first image feature of the image data and the target prompt feature, where the target prompt feature is obtained according to the public opinion event knowledge base of the target domain matching the to-be-processed public opinion data and represents the image semantic guidance information of the target domain; A second determination module for determining the public opinion event category corresponding to the to-be-processed public opinion data according to the comprehensive text feature and the comprehensive image feature.
10. An electronic device, characterized in that, including: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores one or more computer programs executable by the at least one processor, and the one or more computer programs are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-8.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program, when executed by a processor, implements the method according to any one of claims 1-8.
12. A computer program product, characterized in that, Comprising computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, when the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the method according to any one of claims 1-8.
Citation Information
Cited By
Intelligent traffic management public opinion analysis and early warning system based on VIT multi-mode
CN121836697A