False information identification method and device, electronic equipment and storage medium
The combination of search enhancement and large language model through RAG technology solves the lag problem of existing false information recognition software when identifying new types of false information, achieves higher recognition accuracy and simpler operations, and is suitable for the elderly.
Patent Information
- Application Number
- CN202411865075.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-05-16
AI Technical Summary
Existing false information identification software has lag in identifying new types of false information, making it difficult to cover the elderly population, and the operation is complicated and not suitable for use by the elderly.
By obtaining dialogue information, searching enhancements are performed based on RAG technology, the latest false information and cases related to dialogue information are obtained, high-quality prompt information is generated, and input it into a large language model to output recognition results.
It effectively improves the accuracy of the identification of new types of false information by large models, simplifies operations, is suitable for use by the elderly, and improves the security of false information identification software.
Smart Images

Figure CN120011493A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to technical fields such as deep learning and large language models, and especially to methods, devices, electronic devices and storage media for identifying false information. Background Art
[0002] The current method of identifying false information is mainly to have propagandists publish different types of false information on public platforms and regularly update the latest false information. Users can learn and understand the manifestations of false information through relevant platforms, and judge false information based on the learned knowledge, so as to avoid being defrauded of property by criminals. However, this traditional propaganda method is not effective for the elderly. On the one hand, it is difficult to cover all the elderly through manual propaganda, and the propaganda content is difficult for the elderly to fully understand and absorb, especially for new types of false information, which are difficult for the elderly to identify.
[0003] There are some false information identification software on the market, but due to their complicated operation, they are not suitable for use by the elderly and cannot cover the elderly user group. In addition, the existing false information identification software usually uses static recognition, and the false information stored in the database has a lag, resulting in the software being unable to recognize new types of false information and rhetoric. The recognition of new types of false information has a lag, and the recognition accuracy is low. Summary of the invention
[0004] The present invention provides a false information identification method, device, electronic device and storage medium.
[0005] According to a first aspect of the present disclosure, a false information identification method is provided, comprising:
[0006] Get conversation information;
[0007] Performing a search based on the conversation information to obtain search information related to the conversation information;
[0008] generating prompt information based on the dialogue information and the search information;
[0009] The prompt information is input into a large language model, and a corresponding recognition result is output through the large language model.
[0010] According to a second aspect of the present disclosure, a false information identification device is provided, comprising:
[0011] An acquisition module, configured to acquire conversation information;
[0012] A retrieval module, configured to perform a search based on the conversation information to obtain retrieval information related to the conversation information;
[0013] A generating module, configured to generate prompt information based on the dialogue information and the search information;
[0014] The output module is configured to input the prompt information into the large language model and output the corresponding recognition result through the large language model.
[0015] According to a third aspect of the present disclosure, there is provided an electronic device, including:
[0016] at least one processor; and
[0017] a memory communicatively connected to the at least one processor; wherein,
[0018] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute any method in any of the above technical solutions.
[0019] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any one of the methods described in the above technical solutions.
[0020] According to a fifth aspect of the present disclosure, a computer program product is provided, including a computer program, wherein the computer program implements any one of the methods described in the above technical solutions when executed by a processor.
[0021] The present invention uses retrieval enhancement technology to retrieve information related to the user's search intent in real time, performs retrieval enhancement on the dialogue information input by the user, and generates high-quality prompt information, which can effectively improve the recognition accuracy of large models for new types of false information, thereby improving the security of false information recognition software or systems.
[0022] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.
[0024] Figure 1 is a schematic diagram of the steps of the false information identification method in the embodiment of the present disclosure;
[0025] Figure 2 is a schematic diagram of the steps of another false information identification method in an embodiment of the present disclosure;
[0026] Figure 3is an example diagram of picture information containing scene information in an embodiment of the present disclosure;
[0027] Figure 4 is a principle block diagram of a false information identification device in an embodiment of the present disclosure;
[0028] Figure 5 is a principle block diagram of a fusion module in an embodiment of the present disclosure;
[0029] Figure 6 is a flowchart of a false information identification method in an embodiment of the present disclosure;
[0030] Figure 7 The present invention is a block diagram of an electronic device used to implement the false information identification method of the present invention. DETAILED DESCRIPTION
[0031] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0032] In the prior art, false information identification software usually stores the collected false information in a database in advance. After the user inputs the information to be identified, the information to be identified is matched with the false information in the database to identify whether the information to be identified input by the user is false information. However, since the false information in the database is not updated in real time and usually has a lag, and new types of false information emerge in an endless stream, this results in the false information identification software often being unable to identify new types of false information, resulting in reduced security of the false information identification software or system.
[0033] In view of the above technical problems existing in the prior art, the present invention provides a method for identifying false information. Figure 1 As shown, including:
[0034] Step S101, obtaining conversation information.
[0035] Specifically, in this embodiment, a unified input interface can be set by the agent to receive the dialogue information input by the user, and the dialogue information of the user can be obtained through the dialogue of the agent. The form of the dialogue information is not limited, and it can be one of the information such as text information, picture information, voice information, etc., or multi-modal information composed of two or more information.
[0036] Step S102: searching based on the conversation information to obtain search information related to the conversation information.
[0037] Specifically, in this embodiment, RAG (Retrieval-augmented Generation) technology can be used to search the external knowledge base to find the latest false information related to the dialogue information and related cases of false information as retrieval information. For example, the user inputs "I received a text message, please help me check if there is any problem with this text message. The content of the text message is: Dear user, your mobile phone number was selected as a lucky off-site user by the xx column of xx Satellite TV. You will receive 10,000 yuan in cash. Please click to jump to the xx event website to receive it." The intelligent agent can retrieve cases related to the lottery of xx Satellite TV and xx column based on the user's dialogue information. Retrieval is the first step in the RAG process, which retrieves information related to the problem from a pre-established knowledge base. The purpose of this step is to provide useful context information and knowledge support for the subsequent LLM (Large Language Model) generation process.
[0038] Step S103: generating prompt information based on the dialogue information and the search information.
[0039] Step S103 is an enhancement step in the RAG process, which uses the retrieved information as contextual input for generating LLM to enhance the model's ability to understand and answer specific questions. The purpose of this step is to integrate external knowledge into the large model generation process. After retrieving relevant information, the retrieved information is seamlessly integrated into the user's initial prompt information. This step enriches the user's query with factual data, improves its depth and relevance, and makes the generated text content richer, more accurate, and in line with the user's personalized needs.
[0040] Step S104: input the prompt information into the large language model, and output the corresponding recognition result through the large language model.
[0041] Answer generation is the last step of the RAG process. The purpose of this step is to generate answers that meet user needs in combination with LLM. The generator uses the retrieved information as context input and combines it with the large language model to generate text content. For example, for the user input "I received a text message, please help me check if there is any problem with this text message. The text message content is: Dear user, your mobile phone number was selected as a lucky off-site user by the xx column of xx Satellite TV. You will receive 10,000 yuan in cash. Please click to jump to the xx event website to receive it", after retrieving relevant cases of property loss due to false winning information according to the user's prompt, it is spliced with the user's original conversation information as context, and then input into LLM as a new prompt information. LLM generates the corresponding answer as the recognition result based on the prompt information, such as "This text message looks like a false message. Criminals often impersonate popular column groups and send text messages claiming that users have won prizes to defraud money. I suggest not clicking on the link in the text message, and not providing any personal information or paying any fees. At the same time, you can report this text message to the relevant department to prevent more people from being deceived."
[0042] Traditional generation technology is usually based on fixed language models, which makes it difficult to quickly introduce new information. RAG technology can quickly introduce new information and knowledge by updating external knowledge bases, keeping the generated content up to date and the accuracy of the generated content. In the application scenario of false information identification, the present invention combines RAG technology. In the false information identification process, relevant false information and cases can be retrieved in real time according to user prompts, thereby generating new prompt information, enhancing the prompt words (prompt) input by LLM, and can effectively improve the recognition accuracy of large models for new false information, thereby improving the security of false information identification software.
[0043] As an optional implementation, step S101, obtaining dialogue information includes: obtaining at least one of text information, picture information, and voice information.
[0044] Specifically, the user can only input text information or picture information or voice information. The user can input a variety of different types of information through the unified input interface of the intelligent body, including text information (such as text messages, chat records, etc.), picture information (such as screenshots, flyer photos, etc.), voice information (such as phone recordings), etc., so that users can use different methods to input conversation information. Especially for the elderly group, direct typing is inconvenient, and direct voice input or picture input is more convenient to operate, which can simplify the operation of false information identification software or systems and cover more user groups.
[0045] The unified interface for receiving multimodal information is also more convenient for users to use. For example, when a user asks whether there is anything abnormal in a chat record, he can input a screenshot of the chat record and enter the corresponding question "Is there anything wrong with this chat record?". The user does not need to manually enter the chat text, especially since chat records often contain some pictures and emoticons that cannot be expressed in words. By directly inputting pictures, the information conveyed in the pictures can be obtained more accurately.
[0046] As an optional implementation, Figure 2 As shown, step S102, before searching based on the dialogue information to obtain the search information related to the dialogue information, also includes:
[0047] Step S201, in response to the dialogue information being multimodal information, the dialogue information is merged to obtain initial prompt information.
[0048] Specifically, when the dialogue information input by the user appears in multiple forms, such as text information and picture information, or text information and voice information, or text information, picture information, voice information, or picture information and voice information, it is necessary to merge the different types of information and extract the original prompt information in the dialogue information, so as to facilitate subsequent retrieval based on the original prompt information.
[0049] Step S102, searching based on the conversation information, obtaining search information related to the conversation information including:
[0050] A search is performed based on the initial prompt information to obtain search information.
[0051] Specifically, when the dialogue information input by the user includes text information and voice information, the voice information can be converted into text and merged with the text information directly input by the user, and keywords are extracted as initial prompt information, namely, the initial prompt.
[0052] As an optional implementation, step S201, in response to the dialogue information being multimodal information, fusing the dialogue information to obtain initial prompt information, includes:
[0053] In response to the dialog information including the picture information and the text information, the picture information is converted into second text information.
[0054] The second text information is merged with the text information to obtain initial prompt information.
[0055] Specifically, in this embodiment, if the conversation information input by the user contains image information, and if the image contains text information, such as application scenarios such as flyer photos or text message screenshots, OCR (Optical Character Recognition) technology can be used to extract keywords from the image information. For example, when identifying flyer photos or scans, special attention can be paid to the financial institution information, telephone numbers, and money-related keywords contained in the flyer, and then integrated with the original text information input by the user. By integrating multimodal information, more comprehensive, accurate, and reliable initial prompt information can be obtained.
[0056] As an optional implementation, step S201, in response to the dialogue information being multimodal information, fusing the dialogue information to obtain initial prompt information, includes:
[0057] In response to the dialogue information including voice information and text information, converting the voice information into third text information by using a voice recognition technology;
[0058] The third text information is merged with the text information to obtain initial prompt information.
[0059] Specifically, in this embodiment, if the dialogue information input by the user includes voice information, such as application scenarios such as the user inputting a telephone recording or directly inputting voice in real time through a microphone, ASR (Automatic Speech Recognition) technology can be used to convert the voice information into text, and then merge it with the original text information input by the user, so as to obtain more comprehensive, more accurate and more reliable initial prompt information. When the dialogue information input by the user only includes voice information, it can also be converted into text through ASR technology as the initial prompt information.
[0060] As an optional implementation, step S201, in response to the dialogue information being multimodal information, fusing the dialogue information to obtain initial prompt information, includes:
[0061] In response to the dialog information including the picture information, the picture information is converted into second text information.
[0062] In response to the dialogue information including voice information, the voice information is converted into third text information by using a voice recognition technology.
[0063] The text information, the second text information and the third text information are integrated to obtain initial prompt information.
[0064] When the dialogue information input by the user contains information in multiple modes such as text, voice, and pictures, all the modal information can be converted into text information and then integrated. The text information can be formatted according to rules and marked with tags such as prompt or html as initial prompt information. By converting picture information and voice information into text and then integrating it with the original text information input by the user, more comprehensive, accurate, and reliable initial prompt information can be obtained.
[0065] As an optional implementation manner, in response to the dialogue information containing picture information, converting the picture information into second text information includes:
[0066] The optical recognition technology is used to extract keywords from the text information in the picture information to obtain the second text information.
[0067] In one embodiment, for a picture mainly containing text information, OCR technology can be used to recognize the text information in the picture, thereby obtaining the second text information. OCR technology can quickly and accurately recognize text from an image. It can recognize multiple languages and fonts to meet the needs of different scenarios. For example, when identifying a flyer photo, special attention can be paid to the financial institution information, phone number, and money-related keywords contained in the flyer, and then merged with the original text information entered by the user. When the dialogue information entered by the user only includes voice information, it can also be converted into text through OCR technology as the initial prompt information.
[0068] As an optional implementation manner, in response to the dialogue information containing picture information, converting the picture information into second text information includes:
[0069] The application scenario in the picture information is identified, and corresponding picture description information is generated as the second text information.
[0070] In another embodiment, for a picture mainly containing scene information, it is necessary to understand what the picture is trying to express, for example Figure 3 As shown in the figure, a picture of a false information application scenario is reproduced in a cartoon way. The picture only contains a small amount of text description. The picture information can be used to identify the scene through the picture-to-text model to generate the corresponding picture description information. For example, for Figure 3The picture description information obtained by scene recognition is "The picture shows a warning scene of common false information on a social platform. On the left is a young girl who is concentrating on looking at the screen of her mobile phone, which displays tempting information such as 'Join the group to receive free game skins'. On the right is a man whose mobile phone screen also displays similar tempting information. The man holds a fishing net from behind the phone and reaches out to the girl to take her belongings. This reminds us that these information may be traps that use false information to take money." For this type of scene picture, it is impossible to directly convert it into text through OCR technology, and it is necessary to understand the meaning of the picture. Therefore, different picture recognition methods are adopted in this disclosure for different types of picture information. Combining OCR technology and image-to-text method, it can recognize pictures input by users for different application scenarios, generate high-quality initial prompt information, and more accurately identify the user's question intention, thereby effectively improving the large model's recognition accuracy for false information.
[0071] As an optional implementation manner, in response to the dialog information being multimodal information, fusing the dialog information to obtain initial prompt information includes:
[0072] The dialogue information is input into the pre-trained multimodal fusion model for fusion to obtain the initial prompt information.
[0073] In another embodiment of the present disclosure, when the user's conversation information contains information in two or more modalities, the multimodal information can be directly input into the multimodal fusion model to integrate and simplify the input multimodal information and retain only useful key information.
[0074] As an optional implementation, the prompt information is input into the large language model, and the corresponding recognition result output by the large language model includes:
[0075] The prompt information is input into the large language model, and the corresponding analysis information and corresponding response strategies are output through the large language model.
[0076] Specifically, in the recognition of plain text conversation information, LLM outputs analysis information for the text information input by the user. For example, the user inputs the text "I received a text message, please help me check if there is any problem with this text message. The content of the text message is: Dear user, your mobile phone number was drawn as a lucky off-site user by the xx column of xx Satellite TV. You will receive 10,000 yuan in cash. Please click to jump to the xx activity website to receive it." LLM outputs "This text message looks like a false message. Criminals often impersonate popular column groups and send text messages claiming that users have won prizes in order to defraud money" as analysis information. At the same time, it outputs "I suggest not clicking on the link in the text message, and do not provide any personal information or pay any fees. At the same time, you can report this text message to the relevant departments to prevent more people from being deceived" as a response strategy.
[0077] When the conversation information contains image information such as a photo of a flyer, LLM outputs analysis information for the flyer photo input by the user, "The flyer you received has big problems in terms of content. First, the name of the 'Securities Regulatory Commission' mentioned on the flyer is incorrect. The correct name should be 'Securities Regulatory Commission', but there is an extra word 'management' here, which is an obvious mistake. Secondly, the flyer mentions that refunds will be processed through QQ and requires not to publicize the redemption quota to avoid disrupting the order of repayment. This statement is very suspicious. Formal financial institutions or regulatory authorities usually do not handle refunds through social platforms such as QQ. Based on my past experience, this is very likely false information." At the same time, LLM generates "I suggest that you immediately stop any contact with the flyer and do not provide any personal information or bank account information. At the same time, you can report to the local public security agency or consult with a formal financial regulatory agency to confirm the authenticity of this flyer" as a response strategy.
[0078] Similarly, for voice information, the corresponding text can be first recognized, and then the text can be interpreted to obtain the final analysis information and response strategies. By presenting the analysis information and response strategies in the answers, not only the recognition results can be fed back to the user, but also the response strategies can be provided to help the user deal with false information, reducing the risk of users losing money due to false information.
[0079] The present disclosure also provides a false information identification device 400, such as Figure 4 As shown, including:
[0080] The acquisition module 401 is configured to acquire conversation information.
[0081] Specifically, in this embodiment, a unified input interface can be set by the agent to receive the dialogue information input by the user, and the dialogue information of the user can be obtained through the dialogue of the agent. The form of the dialogue information is not limited, and it can be one of the information such as text information, picture information, voice information, etc., or multi-modal information composed of two or more information.
[0082] The retrieval module 402 is configured to perform retrieval based on the dialogue information to obtain retrieval information related to the dialogue information.
[0083] Specifically, in this embodiment, RAG technology can be used to search external knowledge bases to find the latest false information and cases related to the dialogue information as retrieval information. For example, the user inputs "I received a text message, please help me check if there is any problem with this text message. The content of the text message is: Dear user, your mobile phone number was selected as a lucky off-site user by the xx column of xx Satellite TV. You will receive 10,000 yuan in cash. Please click to jump to the xx event website to receive it." The intelligent agent can retrieve cases related to the lottery of xx Satellite TV and xx column based on the user's dialogue information. Retrieval is the first step in the RAG process, which retrieves information related to the problem from a pre-established knowledge base. The purpose of this step is to provide useful contextual information and knowledge support for the subsequent LLM generation process.
[0084] The generating module 403 is configured to generate prompt information based on the dialog information and the search information.
[0085] The generation module 403 uses the search information as contextual input to the LLM to enhance the model's ability to understand and answer specific questions. The purpose of this step is to integrate external knowledge into the large model generation process. After retrieving relevant information, the search information is seamlessly integrated into the user's initial prompt information. This step enriches the user's query with factual data, improves its depth and relevance, and makes the generated text content richer, more accurate, and in line with the user's personalized needs.
[0086] The output module 404 is configured to input the prompt information into the large language model and output the corresponding recognition result through the large language model.
[0087] The last step of the RAG process is to generate answers that meet user needs in combination with LLM. The output module 404 uses the retrieved information as context input and combines it with the large language model to generate text content. For example, for the user input "I received a text message, please help me check if there is any problem with this text message. The text message content is: Dear user, your mobile phone number was selected as a lucky off-site user by the xx column of xx Satellite TV. You will receive 10,000 yuan in cash. Please click to jump to the xx event website to receive it", after retrieving the relevant case according to the user prompt, it is spliced with the user's original conversation information as context, and then input into LLM as a new prompt information. LLM generates the corresponding answer as the recognition result based on the prompt information, for example, "This text message looks like a false message. Criminals often impersonate popular column groups and send text messages claiming that users have won prizes to defraud money. I suggest not clicking on the link in the text message, and not providing any personal information or paying any fees. At the same time, you can report this text message to the relevant department to prevent more people from being deceived."
[0088] Traditional generation technology is usually based on fixed language models, which makes it difficult to quickly introduce new information. RAG technology can quickly introduce new information and knowledge by updating external knowledge bases, keeping the generated content up to date and the accuracy of the generated content. In the application scenario, the present disclosure combines RAG technology to retrieve relevant false information and cases in real time according to user prompts during the false information identification process, thereby generating new prompt information, enhancing the prompt words (prompt) input by LLM, and can effectively improve the recognition accuracy of large models for new false information, thereby improving the security of false information identification software or systems.
[0089] As an optional implementation, the acquisition module 401 acquires the conversation information including: acquiring at least one of text information, picture information, and voice information.
[0090] Specifically, the user can only input text information or picture information or voice information. The user can input a variety of different types of information through the unified input interface of the intelligent body, including text information (such as text messages, chat records, etc.), picture information (such as screenshots, flyer photos, etc.), voice information (such as phone recordings), etc., so that users can use different methods to input conversation information. Especially for the elderly group, direct typing is inconvenient, and direct voice input or picture input is more convenient to operate, which can simplify the operation of false information identification software or systems and cover more user groups.
[0091] The unified interface for receiving multimodal information is also more convenient for users to use. For example, when a user asks whether there is anything abnormal in a chat record, he can input a screenshot of the chat record and enter the corresponding question "Is there anything wrong with this chat record?". The user does not need to manually enter the chat text, especially since chat records often contain some pictures and emoticons that cannot be expressed in words. By directly inputting pictures, the information conveyed in the pictures can be obtained more accurately.
[0092] As an optional implementation, Figure 5 As shown, it also includes:
[0093] The fusion module 405 is configured to fuse the dialogue information to obtain initial prompt information in response to the dialogue information being multimodal information.
[0094] Specifically, when the dialogue information input by the user appears in multiple forms, such as text information and picture information, or text information and voice information, or text information, picture information, voice information, or picture information and voice information, it is necessary to merge the different types of information and extract the original prompt information in the dialogue information, so as to facilitate subsequent retrieval based on the original prompt information.
[0095] The search module 402 searches based on the conversation information, and obtains search information related to the conversation information including:
[0096] A search is performed based on the initial prompt information to obtain search information.
[0097] Specifically, when the dialogue information input by the user includes text information and voice information, the voice information can be converted into text and merged with the text information directly input by the user, and keywords can be extracted as initial prompt information, i.e., initial prompt. Figure 6 As shown, the retrieval module 402 searches the external knowledge base 406 according to the initial prompt information, and searches for the latest false information and cases related to the dialogue information as retrieval information.
[0098] As an optional implementation, Figure 6 As shown, the fusion module 405 includes:
[0099] The picture recognition unit 405a is configured to convert the picture information into second text information in response to the dialogue information containing the picture information and the text information.
[0100] The first fusion unit is configured to fuse the second text information with the text information to obtain initial prompt information.
[0101] Specifically, in this embodiment, if the conversation information input by the user contains image information, and if the image contains text information, such as application scenarios such as flyer photos or text message screenshots, OCR technology can be used to extract keywords from the image information. For example, when identifying flyer photos or scans, special attention can be paid to the financial institution information, telephone numbers, and money-related keywords contained in the flyer, and then integrated with the original text information input by the user.
[0102] As an optional implementation, Figure 6 As shown, the fusion module 405 also includes:
[0103] The speech recognition unit 405b is configured to convert the speech information into third text information by using speech recognition technology in response to the dialogue information containing speech information and text information.
[0104] The second fusion unit is configured to fuse the third text information with the text information to obtain initial prompt information.
[0105] Specifically, in this embodiment, if the conversation information input by the user includes voice information, such as application scenarios such as the user inputting a phone recording or directly inputting voice in real time through a microphone, the ASR technology can be used to convert the voice information into text, and then merged with the original text information input by the user. When the conversation information input by the user only includes voice information, it can also be converted into text through the ASR technology as the initial prompt information.
[0106] As an optional implementation, Figure 6 As shown, the fusion module 405 includes:
[0107] The picture recognition unit 405a is configured to convert the picture information into second text information in response to the dialogue information containing the picture information.
[0108] The speech recognition unit 405b is configured to convert the speech information into third text information by using speech recognition technology in response to the dialogue information containing the speech information.
[0109] The third fusion unit is configured to fuse the text information, the second text information and the third text information to obtain initial prompt information.
[0110] When the dialog information acquired by the acquisition module 401 includes information in multiple modes such as text, voice, and picture, all the text information can be fused after converting all the modal information into text information through the fusion module 405. The text information can be formatted according to rules and marked with tags such as prompt or html as the initial prompt information.
[0111] As an optional implementation manner, in response to the dialogue information containing the picture information, the picture recognition unit 405a converts the picture information into the second text information, including:
[0112] The optical recognition technology is used to extract keywords from the text information in the picture information to obtain the second text information.
[0113] In one embodiment, for images that mainly contain text information, OCR technology can be used to recognize the text information in the image, thereby obtaining the second text information. For example, when recognizing a flyer photo, special attention can be paid to the financial institution information, phone number, and money-related keywords contained in the flyer, and then merged with the original text information entered by the user. When the conversation information entered by the user only includes voice information, it can also be converted into text through OCR technology as the initial prompt information.
[0114] As an optional implementation manner, in response to the dialogue information containing the picture information, the picture recognition unit 405a converts the picture information into the second text information, including:
[0115] Perform scene recognition on the image information and generate corresponding image description information as the second text information.
[0116] In another embodiment, for a picture mainly containing scene information, it is necessary to understand what the picture is trying to express, for example Figure 3 As shown in the figure, a picture of a false information application scenario is reproduced in a cartoon way. The picture only contains a small amount of text description. The picture information can be used to identify the scene through the picture-to-text model to generate the corresponding picture description information. For example, for Figure 3 The picture description information obtained by scene recognition is "The picture shows a warning scene of common false information on a social platform. On the left is a young girl who is concentrating on looking at the screen of her mobile phone, which displays tempting information such as 'Join the group to receive free game skins'. On the right is a man whose mobile phone screen also displays similar tempting information. The man holds a fishing net from behind the phone and reaches out to the girl to take her belongings. This reminds us that these information may be traps that use false information to take money." For this type of scene picture, it is impossible to directly convert it into text through OCR technology, and it is necessary to understand the meaning of the picture. Therefore, different picture recognition methods are adopted in this disclosure for different types of picture information. Combining OCR technology and image-to-text method, it can recognize pictures input by users for different application scenarios, generate high-quality initial prompt information, and more accurately identify the user's question intention, thereby effectively improving the large model's recognition accuracy for false information.
[0117] As an optional implementation, in response to the dialogue information being multimodal information, the fusion module 405 fuses the dialogue information to obtain initial prompt information, including:
[0118] The dialogue information is input into the pre-trained multimodal fusion model for fusion to obtain the initial prompt information.
[0119] In another embodiment of the present disclosure, when the user's conversation information contains information of more than two modalities, the conversation information can be directly input into a multimodal fusion model to integrate and simplify the input multimodal information and retain only useful key information.
[0120] As an optional implementation, the output module 404 inputs the prompt information into the large language model, and outputs the corresponding recognition result through the large language model, including:
[0121] The prompt information is input into the large language model, and the corresponding analysis information and corresponding response strategies are output through the large language model.
[0122] Specifically, in the recognition of plain text conversation information, LLM outputs analysis information for the text information input by the user. For example, the user inputs the text "I received a text message, please help me check if there is any problem with this text message. The content of the text message is: Dear user, your mobile phone number was drawn as a lucky off-site user by the xx column of xx Satellite TV. You will receive 10,000 yuan in cash. Please click to jump to the xx activity website to receive it." LLM outputs "This text message looks like a false message. Criminals often impersonate popular column groups and send text messages claiming that users have won prizes in order to defraud money" as analysis information. At the same time, it outputs "I suggest not clicking on the link in the text message, and do not provide any personal information or pay any fees. At the same time, you can report this text message to the relevant departments to prevent more people from being deceived" as a response strategy.
[0123] When the conversation information contains image information such as a photo of a flyer, LLM outputs analysis information for the flyer photo input by the user, "The flyer you received has big problems in terms of content. First, the name of the 'Securities Regulatory Commission' mentioned on the flyer is incorrect. The correct name should be 'Securities Regulatory Commission', but there is an extra word 'management' here, which is an obvious mistake. Secondly, the flyer mentions that refunds will be processed through QQ and requires not to publicize the redemption quota to avoid disrupting the order of repayment. This statement is very suspicious. Formal financial institutions or regulatory authorities usually do not handle refunds through social platforms such as QQ. Based on my past experience, this is very likely false information." At the same time, LLM generates "I suggest that you immediately stop any contact with the flyer and do not provide any personal information or bank account information. At the same time, you can report to the local public security agency or consult with a formal financial regulatory agency to confirm the authenticity of this flyer" as a response strategy.
[0124] Similarly, for voice information, the corresponding text can be first recognized, and the text can be interpreted to obtain the final analysis information and response strategies. By presenting the analysis information and response strategies in the answers, not only the recognition results are fed back to the user, but also the response strategies can be provided to help the user deal with the false information received, reducing the risk of the user being defrauded of money.
[0125] In the technical solution disclosed herein, the acquisition, storage and application of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0126] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.
[0127] Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0128] like Figure 7As shown, the device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0129] A number of components in the device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0130] The computing unit 701 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 701 performs the various methods and processes described above, such as the false information identification method. For example, in some embodiments, the false information identification method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the false information identification method described above may be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to perform the false information identification method in any other appropriate manner (e.g., by means of firmware).
[0131] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0132] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0133] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0134] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0135] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0136] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0137] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of this disclosure can be achieved, and this document is not limited here.
[0138] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A method for identifying false information, comprising: Get conversation information; Performing a search based on the conversation information to obtain search information related to the conversation information; generating prompt information based on the dialogue information and the search information; The prompt information is input into a large language model, and a corresponding recognition result is output through the large language model.
2. The method according to claim 1, wherein: The acquiring of the conversation information includes: acquiring at least one of text information, picture information, and voice information.
3. The method according to claim 2, wherein: Before performing the search based on the conversation information to obtain the search information related to the conversation information, the method further includes: In response to the dialogue information being multimodal information, fusing the dialogue information to obtain initial prompt information; The searching based on the conversation information to obtain the search information related to the conversation information includes: A search is performed based on the initial prompt information to obtain the search information.
4. The method according to claim 3, wherein: In response to the dialogue information being multimodal information, fusing the dialogue information to obtain initial prompt information includes: In response to the dialog information including the picture information and the text information, converting the picture information into second text information; The second text information is merged with the text information to obtain the initial prompt information.
5. The method according to claim 3, wherein: In response to the dialogue information being multimodal information, fusing the dialogue information to obtain initial prompt information includes: In response to the dialogue information including the voice information and the text information, converting the voice information into third text information by using a voice recognition technology; The third text information is merged with the text information to obtain the initial prompt information.
6. The method according to claim 3, wherein: In response to the dialogue information being multimodal information, fusing the dialogue information to obtain initial prompt information includes: In response to the dialog information including the picture information, converting the picture information into second text information; In response to the dialogue information including the voice information, converting the voice information into third text information by using a voice recognition technology; The text information, the second text information and the third text information are merged to obtain the initial prompt information.
7. The method according to claim 4 or 6, wherein: In response to the dialogue information including the picture information, converting the picture information into the second text information comprises: Keywords are extracted from the text information in the picture information through optical recognition technology to obtain second text information.
8. The method according to claim 4 or 6, wherein: In response to the dialogue information including the picture information, converting the picture information into the second text information comprises: Perform scene recognition on the image information and generate corresponding image description information as the second text information.
9. The method according to claim 3, wherein: In response to the dialogue information being multimodal information, fusing the dialogue information to obtain initial prompt information includes: The dialogue information is input into a pre-trained multimodal fusion model for fusion to obtain the initial prompt information.
10. The method according to any one of claims 1 to 9, wherein: The inputting the prompt information into a large language model and outputting a corresponding recognition result through the large language model comprises: The prompt information is input into a large language model, and corresponding analysis information and a corresponding coping strategy are output through the large language model.
11. A false information identification device, comprising: An acquisition module, configured to acquire conversation information; A retrieval module, configured to perform a search based on the conversation information to obtain retrieval information related to the conversation information; A generating module, configured to generate prompt information based on the dialogue information and the search information; The output module is configured to input the prompt information into the large language model and output the corresponding recognition result through the large language model.
12. The device according to claim 11, wherein The acquisition module acquires the conversation information including: acquiring at least one of text information, picture information, and voice information.
13. The device according to claim 12, wherein: Also includes: a fusion module, configured to, in response to the dialogue information being multimodal information, fuse the dialogue information to obtain initial prompt information; The retrieval module performs retrieval based on the conversation information, and obtains retrieval information related to the conversation information including: A search is performed based on the initial prompt information to obtain the search information.
14. The device according to claim 13, wherein: The fusion module includes: A picture recognition unit, configured to convert the picture information into second text information in response to the dialogue information containing the picture information and the text information; The first fusion unit is configured to fuse the second text information with the text information to obtain the initial prompt information.
15. The device according to claim 13, wherein: The fusion module includes: a speech recognition unit, configured to convert the speech information into third text information by speech recognition technology in response to the dialogue information containing the speech information and the text information; The second fusion unit is configured to fuse the third text information with the text information to obtain the initial prompt information.
16. The device according to claim 13, wherein: The fusion module includes: A picture recognition unit, configured to convert the picture information into second text information in response to the dialogue information containing the picture information; a speech recognition unit configured to convert the speech information into third text information by speech recognition technology in response to the dialogue information containing the speech information; The third fusion unit is configured to fuse the text information, the second text information and the third text information to obtain the initial prompt information.
17. The device according to claim 14 or 16, wherein: The picture recognition unit converts the picture information into second text information in response to the dialogue information containing the picture information, including: Keywords are extracted from the text information in the picture information through optical recognition technology to obtain second text information.
18. The device according to claim 14 or 16, wherein: The picture recognition unit converts the picture information into second text information in response to the dialogue information containing the picture information, including: Perform scene recognition on the image information and generate corresponding image description information as the second text information.
19. The device according to claim 13, wherein: In response to the dialogue information being multimodal information, the fusion module fuses the dialogue information to obtain initial prompt information, including: The dialogue information is input into a pre-trained multimodal fusion model for fusion to obtain the initial prompt information.
20. The device according to any one of claims 11 to 19, wherein: The output module inputs the prompt information into the large language model, and outputs the corresponding recognition result through the large language model, including: The prompt information is input into a large language model, and corresponding analysis information and a corresponding coping strategy are output through the large language model.
21. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 10.
22. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-10.
23. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 10.