A social media sensitive information identification method, device, equipment and storage medium
By identifying and filtering sensitive information in social media, combining user profiles to determine the reliability of the information, and using generative adversarial networks to train models, the problem of interference from fake information in social media has been solved, improving the credibility and analysis efficiency of sensitive information.
Patent Information
- Application Number
- CN202310231833.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-09
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2043-03-09
AI Technical Summary
The information filtering process on social media is susceptible to interference from fabricated information, leading to inaccurate analysis results, which could have serious consequences, especially in the military field.
Sensitive information in social media is identified by pre-set sensitive information identification technology and fake information identification model. The reliability of the information is judged by combining user profile identification model. Generative adversarial network is used to train the sensitive information identification model, filter out non-fake sensitive information and generate individual behavior profiles to determine the final sensitive information.
It has improved the credibility of sensitive information on social media, enhanced the accuracy and efficiency of information analysis, and reduced interference from fake information.
Smart Images

Figure CN116226415B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of information recognition, and in particular relates to a social media sensitive information recognition method and device, equipment and a storage medium. BACKGROUND
[0002] In recent years, the amount of video image data in social media has rapidly increased, and by analyzing and predicting the content such as pictures and videos on social media, a lot of valuable information can be obtained. However, in the process of screening information in social media, it is easy to be disturbed by the fake information therein, which leads to inaccurate results of the final analysis, and in some fields, such as the military field, it may cause serious consequences. Therefore, how to identify the reliability of sensitive information in social media is a problem to be solved in the field. SUMMARY
[0003] Therefore, the purpose of the present application is to provide a social media sensitive information recognition method, device, equipment and storage medium, which can improve the reliability of sensitive information by combining user portraits. The specific scheme is as follows:
[0004] In the first aspect, the present application provides a social media sensitive information recognition method, comprising:
[0005] grabbing the content sent by the user on the social media to obtain to-be-recognized information;
[0006] identifying sensitive information in the to-be-recognized information by a preset sensitive information recognition technology to obtain first sensitive information, and determining non-fake sensitive information from the first sensitive information based on a preset fake information recognition model to obtain second sensitive information;
[0007] generating an individual behavior portrait of the corresponding user according to the second sensitive information, and identifying a target portrait of a preset reliable individual from the individual behavior portrait by using a preset user portrait recognition model;
[0008] determining the second sensitive information corresponding to the target portrait as the final sensitive information.
[0009] Optionally, the judgment of the to-be-recognized information by the preset sensitive information recognition technology to obtain the first sensitive information comprises:
[0010] screening the to-be-recognized information according to a preset sensitive keyword set to obtain to-be-recognized information containing sensitive keywords;
[0011] judging the image and the semantics in the to-be-recognized information by a preset image recognition technology and a preset semantic recognition technology to obtain to-be-recognized information containing sensitive images and / or sensitive semantics;
[0012] obtaining the first sensitive information based on the to-be-identified information containing the sensitive keyword and the to-be-identified information containing the sensitive image and / or the sensitive semantics.
[0013] Optionally, the determining, by the preset fake information identification model, that the first sensitive information is non-fake information as the second sensitive information, comprises:
[0014] judging, based on a preset multimedia file fake identification model, the first sensitive information to obtain multimedia file fake sensitive information;
[0015] judging, based on a preset semantic modification identification model, the first sensitive information to obtain semantic modification sensitive information;
[0016] obtaining the second sensitive information based on the multimedia file fake sensitive information and the semantic modification sensitive information.
[0017] Optionally, before the judging, based on the preset multimedia file fake identification model, the first sensitive information to obtain multimedia file fake sensitive information, the method further comprises:
[0018] generating, by a generator of a generative adversarial network, a new multimedia vector sequence from an initial multimedia file sent by an existing natural person user;
[0019] iteratively training, by a discriminator of the generative adversarial network, the new multimedia vector sequence and an initial multimedia vector sequence corresponding to the initial multimedia file, to obtain the preset multimedia file fake identification model for identifying whether a multimedia file corresponding to a multimedia vector sequence is a file generated out of nothing, so as to judge, based on the preset multimedia file fake identification model, the first sensitive information to obtain multimedia file fake sensitive information.
[0020] Optionally, before the judging, based on the preset semantic modification identification model, the first sensitive information to obtain semantic modification sensitive information, the method further comprises:
[0021] performing semantic-level modification on an initial multimedia file sent by an existing natural person user to obtain a modified multimedia vector sequence;
[0022] generating, by a generator of a generative adversarial network, a new modified multimedia vector sequence according to the modified multimedia vector sequence, and iteratively training, by a discriminator, the new modified multimedia vector sequence and the modified multimedia vector sequence to obtain the semantic modification identification model for identifying whether a multimedia file corresponding to a multimedia vector sequence is a semantic modification file, so as to judge, based on the preset semantic modification identification model, the first sensitive information to obtain semantic modification sensitive information.
[0023] Optionally, before the generating individual behavior portraits of corresponding users according to the second sensitive information and identifying a target portrait of a preset reliable individual from the individual behavior portraits by using a preset user portrait recognition model, the method further comprises:
[0024] collecting existing natural person user information, and generating original individual behavior portraits according to the natural person user information, so as to obtain a corresponding original user information vector sequence according to the original individual behavior portraits;
[0025] generating a new user information vector sequence according to the original user information vector sequence by using a generator of a generative adversarial network, and iteratively training the new user information vector sequence and the original user information vector sequence by using a discriminator to obtain the preset user portrait recognition model for identifying individual behavior portraits, so as to identify the target portrait of the preset reliable individual from the individual behavior portraits by using the preset user portrait recognition model.
[0026] Optionally, after the determining the final sensitive information by using the second sensitive information corresponding to the target portrait, the method further comprises:
[0027] respectively saving the final sensitive information of the individual behavior portrait as the target portrait and the sensitive information of the individual behavior portrait as non-target portrait, so as to be processed by a staff.
[0028] In a second aspect, the present application provides a social media sensitive information recognition device, comprising:
[0029] an information grabbing module, configured to grab content sent by a user on a social media to obtain to-be-identified information;
[0030] a sensitive information recognition module, configured to identify sensitive information in the to-be-identified information by using a preset sensitive information recognition technology to obtain first sensitive information, and determine non-fake sensitive information from the first sensitive information based on a preset fake information recognition model to obtain second sensitive information;
[0031] a portrait recognition module, configured to generate individual behavior portraits of corresponding users according to the second sensitive information, and identify a target portrait of a preset reliable individual from the individual behavior portraits by using a preset user portrait recognition model;
[0032] a sensitive information determination module, configured to determine final sensitive information by using the second sensitive information corresponding to the target portrait.
[0033] In a third aspect, the present application provides an electronic device, comprising:
[0034] a memory, configured to save a computer program;
[0035] A processor is configured to execute the computer program to implement the social media sensitive information identification method as described above.
[0036] In a fourth aspect, the present application provides a computer readable storage medium for storing a computer program, which, when executed by a processor, implements the social media sensitive information identification method as described above.
[0037] It can be seen that the present application can capture the content sent by the user on the social media to obtain the to-be-identified information, identify the sensitive information in the to-be-identified information by using a preset sensitive information identification technology to obtain first sensitive information, determine the non-fake sensitive information from the first sensitive information based on a preset fake information identification model to obtain second sensitive information, generate the individual behavior portrait of the corresponding user according to the second sensitive information, and identify the target portrait of the preset reliable individual from the individual behavior portrait by using a preset user portrait identification model, and then determine the final sensitive information corresponding to the second sensitive information of the target portrait. It can be seen that the present application can determine the final sensitive information by combining the individual behavior portrait corresponding to the sensitive information after determining the sensitive information by using the preset sensitive information identification technology and the preset fake information identification model, so as to improve the credibility of the final sensitive information, thereby facilitating the subsequent sensitive information analysis operation. BRIEF DESCRIPTION OF DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only embodiments of the present application, and those skilled in the art can obtain other drawings according to the provided drawings without creative labor.
[0039] Figure 1 A social media sensitive information identification method flow chart disclosed by the present application;
[0040] Figure 2 A specific social media sensitive information identification method flow chart disclosed by the present application;
[0041] Figure 3 Another social media sensitive information identification method flow chart disclosed by the present application;
[0042] Figure 4 A social media sensitive information identification device structure schematic diagram disclosed by the present application;
[0043] Figure 5 A structure diagram of an electronic device disclosed by the present application. DETAILED DESCRIPTION
[0044] With reference to the drawings of the embodiments of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments of the present application, all the other embodiments obtained by a person of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0045] Image modification is a special task in the field of image generation, which requires generating a picture after modifying the original picture. At present, most of the picture manipulation and generation are carried out at the pixel level. With the advancement of technology, semantic-based image modification and generation become possible. Further, as a kind of deep learning model, the generative adversarial network (GAN) has the ability to continuously improve modeling under the game. By training the discriminator and the generator alternately, it can be made to be against each other, so as to finally realize the image generation ability of making the fake real and the identification ability of the fake image. It can be applied to the analysis and fake identification process of social media resources (especially picture and video resources). Therefore, the present application can effectively judge the reliability of information in social media according to the generative adversarial network.
[0046] Referring to Figure 1 The embodiments of the present application disclose a social media sensitive information identification method, which comprises the following steps:
[0047] Step S11, grabbing the content sent by the user on the social media to obtain the to-be-identified information.
[0048] It can be understood that the technical solutions of the present application are aimed at the content sent by the user in the social media, and the content sent by the user can be grabbed from the social media as the to-be-identified information. In specific embodiments, it can include grabbing the text information, picture information, video information and the like sent by the user.
[0049] Step S12, identifying the sensitive information in the to-be-identified information by a preset sensitive information identification technology to obtain first sensitive information, and determining the non-fake sensitive information from the first sensitive information based on a preset fake information identification model to obtain second sensitive information.
[0050] In the embodiment, the judging the to-be-identified information by the preset sensitive information identification technology to obtain the first sensitive information can include: screening the to-be-identified information according to a preset sensitive keyword set to obtain to-be-identified information containing a sensitive keyword; judging images and semantics in the to-be-identified information by a preset image identification technology and a preset semantic identification technology respectively to obtain to-be-identified information containing a sensitive image and / or a sensitive semantic; and obtaining the first sensitive information based on the to-be-identified information containing the sensitive keyword and the to-be-identified information containing the sensitive image and / or the sensitive semantic. In a specific embodiment, whether the to-be-identified information contains sensitive information can be judged by a preset sensitive keyword set and by image identification, semantic identification and other technologies. Further, to-be-identified information containing a sensitive keyword and / or containing sensitive pictures, semantics and other sensitive information can be determined as the first sensitive information.
[0051] Further, the determining the second sensitive information based on the first sensitive information determined as non-fake information by the preset fake information identification model comprises: determining multimedia file fake sensitive information based on the first sensitive information by a preset multimedia file fake identification model; determining semantic modification sensitive information based on the first sensitive information by a preset semantic modification identification model; and determining the second sensitive information based on the multimedia file fake sensitive information and the semantic modification sensitive information. Specifically, the first sensitive information can be determined as information generated out of thin air or information modified from existing materials. In this embodiment, the fake information can be determined by a pre-trained model. It can be understood that before the first sensitive information is determined as multimedia file fake sensitive information based on the preset multimedia file fake identification model, the method can further comprise: generating a new multimedia vector sequence based on an initial multimedia file sent by an existing natural person user by using a generator of a generative adversarial network; and iteratively training the new multimedia vector sequence and an initial multimedia vector sequence corresponding to the initial multimedia file by using a discriminator of the generative adversarial network to obtain the preset multimedia file fake identification model for identifying whether a multimedia file corresponding to the multimedia vector sequence is a file generated out of thin air, so as to determine the first sensitive information as multimedia file fake sensitive information based on the preset multimedia file fake identification model. In a specific embodiment, multimedia (video, image) sent by an existing natural person user can be collected, and an original multimedia vector sequence can be formed based on the multimedia. The original multimedia vector sequence is input into a generator of a GAN network for training, and a generated multimedia vector sequence is output. The original multimedia vector sequence and the generated multimedia vector sequence are input into a discriminator of the GAN for training and updating parameters in the discriminator, so that the discriminator can distinguish between the original multimedia vector sequence and the generated multimedia vector sequence. The generator and the discriminator are alternately trained in this way, and a multimedia generator and a discriminator with sufficient strength can be obtained. The discriminator can be used to identify whether a multimedia file is a file generated out of thin air.
[0052] Correspondingly, before the preset semantic modification recognition model judges the first sensitive information to obtain semantic modification sensitive information, the method can further include: performing semantic level modification on an initial multimedia file sent by an existing natural person user to obtain a modified multimedia vector sequence; generating a new modified multimedia vector sequence by using a generator of a generative adversarial network according to the modified multimedia vector sequence, and iteratively training the new modified multimedia vector sequence and the modified multimedia vector sequence by using a discriminator to obtain the semantic modification recognition model for identifying whether a multimedia file corresponding to the multimedia vector sequence is a semantic modification file, so as to judge the first sensitive information based on the preset semantic modification recognition model to obtain semantic modification sensitive information. In a specific embodiment, a tweet (video, image) sent by an existing natural person user is collected, the content in the original tweet is modified at a semantic level to form a modified information vector sequence; then the modified information vector sequence is input into a generator of a GAN network for training, and a generated information vector sequence is output; the modified information vector sequence and the generated information vector sequence are input into a discriminator of the GAN, and the parameters in the discriminator are trained and updated so that the discriminator can distinguish the modified information vector sequence from the generated information vector sequence; the generator and the discriminator are alternately trained, and a multimedia semantic level modification generator and discriminator with sufficient strength can be obtained. The discriminator can be used to identify whether a multimedia file is a file obtained by modifying another original file at a semantic level.
[0053] In the present application, the aforementioned preset multimedia file forgery recognition model and the preset semantic modification recognition model can filter out multimedia files that are not generated out of thin air and not modified from the first sensitive information. These multimedia files have relatively high authenticity and can be confirmed as the second sensitive information.
[0054] Step S13, generating an individual behavior portrait of a corresponding user according to the second sensitive information, and identifying a target portrait of a preset reliable individual from the individual behavior portrait by using a preset user portrait recognition model.
[0055] In the embodiments of the present application, after obtaining the second sensitive information based on the generative adversarial network, the individual behavior portrait corresponding to the user can be further generated according to the second sensitive information. Specifically, the individual behavior portrait of the user corresponding to the second sensitive information can be generated, which is used to determine whether the user of the information is a normal natural person or a robot. Further, in specific embodiments, according to the sensitive information, the individual behavior portrait of the user can be generated according to other information sent by the user corresponding to the sensitive information. It should be pointed out that if there is no corresponding relationship between the individual behavior portrait and the sensitive information, for example, the portrait of the user is a doctor, and the sensitive information sent is related to the military, it is considered that the individual behavior portrait does not correspond to the sensitive information. At this time, the credibility of the sensitive information can be appropriately reduced. In the embodiments of the present application, the preset user portrait recognition model can be used to identify the target portrait and the corresponding sensitive information from the individual behavior portrait and the corresponding second sensitive information.
[0056] In step S14, the second sensitive information corresponding to the target portrait is determined as the final sensitive information.
[0057] In the embodiments, after the target portrait is identified according to the preset user portrait recognition model, the second sensitive information corresponding to the target portrait can be determined as the final sensitive information, and at this time, the credibility of the final sensitive information is relatively high. It should be pointed out that after the second sensitive information corresponding to the target portrait is determined as the final sensitive information, the final sensitive information of the individual behavior portrait as the target portrait and the sensitive information of the individual behavior portrait as the non-target portrait can be saved respectively for the staff to process. It can be understood that through the foregoing steps, the final sensitive information with relatively high credibility can be obtained, and the sensitive information with relatively low credibility or even untrustworthy can be screened out. In specific embodiments, the sensitive information in the process can be classified and stored according to different credibility of the sensitive information, so that the staff can analyze the sensitive information in depth in the future.
[0058] As can be seen, in the present application, a plurality of sensitive information recognition models can be trained by using the generative adversarial network, and the authenticity of the sensitive information can be detected from different directions. Moreover, the credibility of the sensitive information can be judged in combination with the individual behavior portrait of the user corresponding to the sensitive information. If the individual behavior portrait corresponding to the sensitive information is a non-robot, the credibility of the sensitive information can be appropriately improved. Then, the final sensitive information is classified and stored according to the credibility. In this way, the present application can improve the credibility of the final sensitive information, and further improve the efficiency of obtaining valuable information from social media.
[0059] Referring to Figure 2As shown, the embodiment of the present application discloses a social media sensitive information recognition method, comprising:
[0060] Step S21, grabbing the content sent by the user on the social media to obtain the to-be-recognized information.
[0061] Step S22, recognizing the sensitive information in the to-be-recognized information by a preset sensitive information recognition technology to obtain first sensitive information, and determining non-fake sensitive information from the first sensitive information based on a preset fake information recognition model to obtain second sensitive information.
[0062] Step S23, collecting existing natural person user information, and generating an original individual behavior portrait according to the natural person user information to obtain a corresponding original user information vector sequence according to the original individual behavior portrait.
[0063] In the embodiment, it can be understood that before the individual behavior portrait of the corresponding user is generated according to the second sensitive information, and the target portrait of the preset reliable individual is recognized from the individual behavior portrait by using the preset user portrait recognition model, the method can further comprise: collecting existing natural person user information, and generating an original individual behavior portrait according to the natural person user information to obtain a corresponding original user information vector sequence according to the original individual behavior portrait. Specifically, the information of the existing natural person user is first collected, the corresponding original individual behavior portrait is generated according to the information, and further, the original user information vector sequence for inputting the GAN generator can be generated by using the original individual behavior portrait.
[0064] Step S24, generating a new user information vector sequence according to the original user information vector sequence by using a generator of a generative adversarial network, and iteratively training the new user information vector sequence and the original user information vector sequence by using a discriminator to obtain the preset user portrait recognition model for recognizing the individual behavior portrait, so as to recognize the target portrait of the preset reliable individual from the individual behavior portrait by using the preset user portrait recognition model.
[0065] Further, after obtaining the original user information vector sequence, the generator can be used to generate a new user information vector sequence according to the original user information vector sequence; then the new user information vector sequence and the original user information vector sequence are trained by the discriminator, so that a better user information generator and discriminator can be obtained. It can be understood that the discriminator can be used to judge the possibility that the user corresponding to the sensitive information is a normal natural person or a non-robot.
[0066] In specific embodiments, as Figure 3As shown, the tweet content sent by the user on the social media can be judged in combination with the individual behavior portrait, the credibility of the tweet is determined by calculating the individual behavior portrait of the tweet-sending user, and the final sensitive information that is reliable is screened out and handed over to the information analyst.
[0067] Step S25, generating an individual behavior portrait of the corresponding user according to the second sensitive information, and identifying a target portrait of a preset reliable individual from the individual behavior portrait by using a preset user portrait recognition model.
[0068] Step S26, determining the second sensitive information corresponding to the target portrait as the final sensitive information.
[0069] More specific processing procedures of the above steps S21, S22, S25 and S26 can refer to the corresponding contents disclosed in the foregoing embodiments, which will not be repeated here.
[0070] As can be seen, the embodiments of the present application can train the individual behavior portrait corresponding to the user by using the generator and the discriminator in the generative adversarial network, and can obtain a user portrait detection model for judging whether the user corresponding to the information is likely to be a robot; in this way, the present application can evaluate the credibility of the sensitive information in combination with the individual behavior portrait corresponding to the information, which can improve the accuracy of sensitive information recognition and further improve the efficiency of subsequent sensitive information analysis.
[0071] As Figure 4 The social media sensitive information recognition device disclosed by the present application is shown, which comprises:
[0072] The information grabbing module 11 is configured to grab the content sent by the user on the social media to obtain the to-be-recognized information.
[0073] The sensitive information recognition module 12 is configured to recognize the sensitive information in the to-be-recognized information by using a preset sensitive information recognition technology to obtain the first sensitive information, and determine the non-fake sensitive information from the first sensitive information based on a preset fake information recognition model to obtain the second sensitive information.
[0074] The portrait recognition module 13 is configured to generate an individual behavior portrait of the corresponding user according to the second sensitive information, and identify a target portrait of a preset reliable individual from the individual behavior portrait by using a preset user portrait recognition model.
[0075] The sensitive information determination module 14 is configured to determine the second sensitive information corresponding to the target portrait as the final sensitive information.
[0076] Therefore, the application can determine the final sensitive information by combining the individual behavior portrait corresponding to the sensitive information after judging the sensitive information through the preset sensitive information identification technology and the preset fake information identification model, can improve the credibility of the final sensitive information, and can improve the efficiency of obtaining valuable information in subsequent sensitive information analysis operations.
[0077] In a specific embodiment, the sensitive information identification module 12 can include:
[0078] A keyword screening unit is configured to screen the to-be-identified information according to a preset sensitive keyword set to obtain to-be-identified information containing sensitive keywords.
[0079] A sensitive information identification unit is configured to judge images and semantics in the to-be-identified information through a preset image identification technology and a preset semantic identification technology to obtain to-be-identified information containing sensitive images and / or sensitive semantics.
[0080] A first sensitive information determination unit is configured to obtain the first sensitive information based on the to-be-identified information containing sensitive keywords and the to-be-identified information containing sensitive images and / or sensitive semantics.
[0081] In a specific embodiment, the sensitive information identification module 12 can include:
[0082] A first fake information identification sub-module is configured to judge the first sensitive information based on a preset multimedia file fake identification model to obtain multimedia file fake sensitive information.
[0083] A second fake information identification sub-module is configured to judge the first sensitive information based on a preset semantic modification identification model to obtain semantic modification sensitive information.
[0084] A second sensitive information determination unit is configured to obtain the second sensitive information based on the multimedia file fake sensitive information and the semantic modification sensitive information.
[0085] In another specific embodiment, the first fake information identification sub-module can further include:
[0086] A first vector sequence generation unit is configured to generate a new multimedia vector sequence according to an existing natural person user sent initial multimedia file by using a generator of a generative adversarial network.
[0087] The first forgery identification model training unit is configured to train the new multimedia vector sequence and the initial multimedia vector sequence corresponding to the initial multimedia file by using a discriminator of a generative adversarial network to obtain the preset multimedia file forgery identification model for identifying whether the multimedia file corresponding to the multimedia vector sequence is a file generated out of nothing, so as to judge the first sensitive information based on the preset multimedia file forgery identification model to obtain multimedia file forgery sensitive information.
[0088] In another specific embodiment, the second forgery information identification sub-module can further include:
[0089] The second vector sequence generation unit is configured to perform semantic-level modification on an initial multimedia file sent by an existing natural person user to obtain a modified multimedia vector sequence.
[0090] The second forgery identification model training unit is configured to generate a new modified multimedia vector sequence by using a generator of a generative adversarial network according to the modified multimedia vector sequence, and train the new modified multimedia vector sequence and the modified multimedia vector sequence by using a discriminator to obtain the semantic modification identification model for identifying whether the multimedia file corresponding to the multimedia vector sequence is a semantic modification file, so as to judge the first sensitive information based on the preset semantic modification identification model to obtain semantic modification sensitive information.
[0091] In a specific embodiment, the portrait identification module 13 can further include:
[0092] The third vector sequence generation unit is configured to collect existing natural person user information, and generate an original individual behavior portrait according to the natural person user information to obtain a corresponding original user information vector sequence according to the original individual behavior portrait.
[0093] The portrait identification model training unit is configured to generate a new user information vector sequence by using a generator of a generative adversarial network according to the original user information vector sequence, and train the new user information vector sequence and the original user information vector sequence by using a discriminator to obtain the preset user portrait identification model for identifying an individual behavior portrait, so as to identify a target portrait of a preset reliable individual from the individual behavior portrait by using the preset user portrait identification model.
[0094] In a specific embodiment, the sensitive information determination module 14 can further include:
[0095] The information saving unit is configured to save the final sensitive information of the individual behavior portrait as the target portrait and the sensitive information of the individual behavior portrait as the non-target portrait, respectively, so as to be processed by a staff.
[0096] Further, the embodiment of the present application further discloses an electronic device, Figure 5 is an electronic device 20 structure diagram shown according to an exemplary embodiment, the contents in the figure cannot be considered as any limitation on the scope of use of the present application.
[0097] Figure 5 An electronic device 20 structure diagram provided by the embodiment of the present application. The electronic device 20, specifically can include: at least one processor 21, at least one memory 22, power supply 23, communication interface 24, input output interface 25 and communication bus 26. Wherein, the memory 22 is used for storing computer programs, the computer programs are loaded and executed by the processor 21, to realize the related steps in the social media sensitive information identification method disclosed by any of the preceding embodiments. In addition, the electronic device 20 in the embodiment of the present application can be an electronic computer.
[0098] In the embodiment, the power supply 23 is used for providing working voltage for each hardware device on the electronic device 20; the communication interface 24 can create data transmission channel between the electronic device 20 and the external device, and the communication protocol followed by the communication interface 24 can be any communication protocol applicable to the technical solution of the present application, which is not limited specifically herein; the input output interface 25 is used for obtaining external input data or outputting data to the outside world, and the specific interface type can be selected according to the specific application needs, which is not limited specifically herein.
[0099] In addition, the memory 22 as the carrier of resource storage can be read-only memory, random access memory, disk or optical disk, etc., and the resources stored thereon can include operating system 221, computer program 222, etc., and the storage mode can be temporary storage or permanent storage.
[0100] Wherein, the operating system 221 is used for managing and controlling each hardware device on the electronic device 20 and the computer program 222, which can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program capable of completing the social media sensitive information identification method executed by the electronic device 20 disclosed by any of the preceding embodiments, the computer program 222 can further include computer programs capable of completing other specific work.
[0101] Further, the present application further discloses a computer readable storage medium for storing computer programs; wherein the computer programs are executed by the processor to realize the social media sensitive information identification method disclosed above. The specific steps of the method can refer to the corresponding contents disclosed in the preceding embodiments, which will not be repeated here.
[0102] The various embodiments described in the specification are progressive in nature, and each embodiment highlights the differences from other embodiments. The same or similar parts among the various embodiments can be mutually referred to. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method part.
[0103] Those skilled in the art will further appreciate that the individual steps of the examples described in connection with the embodiments disclosed herein can be embodied in electronic hardware, computer software, or combinations of both. The various examples have been described in relation to the described embodiments, as a means of generalizing the interchangeability of hardware and software features, and distinguishing the features from other examples, which are further within the scope of other examples. The described features are implemented in hardware and software in different combinations based on the specific application and design constraints imposed on the overall architecture.
[0104] The steps of a method or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module can reside in random access memory (RAM), flash memory, read-only memory (ROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0105] Finally, it needs to be pointed out that in this document, the relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of other identical elements in the process, method, article or device including the element.
[0106] The above describes the technical solutions provided by the present application in detail, and the principles and implementation manners of the present application are described by using specific examples. The above description of the embodiments is only to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, the specific implementation manner and application range can be changed; in view of the above, the content of the specification should not be understood as limiting the present application.
Claims
1. A method for social media sensitive information recognition, the method comprising: The method comprises the following steps: grabbing content sent by a user on social media to obtain to-be-identified information; identifying sensitive information in the to-be-identified information by using a preset sensitive information identification technology to obtain first sensitive information, and determining non-fake sensitive information from the first sensitive information based on a preset fake information identification model to obtain second sensitive information; the second sensitive information is information obtained based on multimedia file fake sensitive information and semantic modification sensitive information, the multimedia file fake sensitive information is information obtained by judging the first sensitive information based on a preset multimedia file fake identification model, and the semantic modification sensitive information is information obtained by judging the first sensitive information based on a preset semantic modification identification model; generating an individual behavior portrait of a corresponding user according to the second sensitive information, and identifying a target portrait of a preset reliable individual from the individual behavior portrait by using a preset user portrait identification model; the preset reliable individual represents that the related user is not a robot; determining the second sensitive information corresponding to the target portrait as final sensitive information; Before the step of generating an individual behavior portrait of a corresponding user according to the second sensitive information, and identifying a target portrait of a preset reliable individual from the individual behavior portrait by using a preset user portrait identification model, the method further comprises the following steps: collecting existing natural person user information, and generating an original individual behavior portrait according to the natural person user information to obtain a corresponding original user information vector sequence; generating a new user information vector sequence according to the original user information vector sequence by using a generator of a generative adversarial network, and iteratively training the new user information vector sequence and the original user information vector sequence by using a discriminator to obtain the preset user portrait identification model for identifying individual behavior portraits, so as to identify a target portrait of a preset reliable individual from the individual behavior portrait by using the preset user portrait identification model.
2. The social media sensitive information recognition method of claim 1, wherein, The step of identifying sensitive information in the to-be-identified information by using a preset sensitive information identification technology to obtain first sensitive information comprises the following steps: screening the to-be-identified information according to a preset sensitive keyword set to obtain to-be-identified information containing sensitive keywords; judging images and semantics in the to-be-identified information by using a preset image identification technology and a preset semantic identification technology to obtain to-be-identified information containing sensitive images and / or sensitive semantics; obtaining the first sensitive information based on the to-be-identified information containing sensitive keywords and the to-be-identified information containing sensitive images and / or sensitive semantics.
3. The social media sensitive information recognition method of claim 1, wherein, Before the step of judging the first sensitive information based on a preset multimedia file fake identification model to obtain multimedia file fake sensitive information, the method further comprises the following steps: generating a new multimedia vector sequence according to an initial multimedia file sent by an existing natural person user by using a generator of a generative adversarial network; The discriminator of the generative adversarial network is used to iteratively train the new multimedia vector sequence and the initial multimedia vector sequence corresponding to the initial multimedia file, to obtain the preset multimedia file forgery identification model for identifying whether the multimedia vector sequence corresponds to a multimedia file that is a file generated out of nothing, so as to judge the first sensitive information based on the preset multimedia file forgery identification model to obtain multimedia file forgery sensitive information.
4. The social media sensitive information recognition method of claim 1, wherein, Before the judging of the first sensitive information based on the preset semantic modification identification model to obtain semantic modification sensitive information, the method further comprises: performing semantic level modification on an initial multimedia file sent by an existing natural person user to obtain a modified multimedia vector sequence; generating a new modified multimedia vector sequence from the modified multimedia vector sequence by using a generator of a generative adversarial network, and iteratively training the new modified multimedia vector sequence and the modified multimedia vector sequence by using a discriminator to obtain the semantic modification identification model for identifying whether the multimedia vector sequence corresponds to a semantic modification file, so as to judge the first sensitive information based on the preset semantic modification identification model to obtain semantic modification sensitive information.
5. The social media sensitive information recognition method of any one of claims 1 to 4, characterized in that, After the second sensitive information corresponding to the target portrait is determined as final sensitive information, the method further comprises: saving the final sensitive information of the individual behavior portrait as a target portrait and the sensitive information of the individual behavior portrait as a non-target portrait respectively, so as to be processed by a staff.
6. A social media sensitive information identification apparatus, characterized by, Comprise: An information grabbing module is configured to grab content sent by a user on a social media to obtain to-be-identified information. A sensitive information identification module is configured to identify sensitive information in the to-be-identified information by using a preset sensitive information identification technology to obtain first sensitive information, and determine non-forged sensitive information from the first sensitive information based on a preset forgery information identification model to obtain second sensitive information. The second sensitive information is obtained based on multimedia file forgery sensitive information and semantic modification sensitive information, the multimedia file forgery sensitive information is obtained based on a judgment of the first sensitive information by using a preset multimedia file forgery identification model, and the semantic modification sensitive information is obtained based on a judgment of the first sensitive information by using a preset semantic modification identification model. A portrait identification module is configured to generate an individual behavior portrait of a corresponding user based on the second sensitive information, and identify a target portrait of a preset reliable individual from the individual behavior portrait by using a preset user portrait identification model. The preset reliable individual indicates that the related user is not a robot. A sensitive information determination module is configured to determine the second sensitive information corresponding to the target portrait as final sensitive information. The portrait identification module further comprises: A third vector sequence generation unit is configured to collect information of an existing natural person user, generate an original individual behavior portrait based on the information of the natural person user, and obtain a corresponding original user information vector sequence based on the original individual behavior portrait. The image recognition model training unit is configured to generate a new user information vector sequence from the original user information vector sequence by using a generator of a generative adversarial network, and to iteratively train the new user information vector sequence and the original user information vector sequence by using a discriminator to obtain the preset user image recognition model for recognizing an individual behavior image, so as to recognize a target image of a preset reliable individual from the individual behavior image by using the preset user image recognition model.
7. An electronic device, comprising: Comprising: a memory for saving a computer program; a processor for executing the computer program to implement the social media sensitive information recognition method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, a memory for saving a computer program, which, when executed by a processor, implements the social media sensitive information recognition method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Automatic text generation method based on class label sequence generative adversarial model
CN111259650A
False news identification method, device, equipment and chip
CN113704400A
Text recognition method and device, computer equipment and computer readable storage medium
CN115659965A
Image recognition method and apparatus, computing device and computer-readable storage medium
WO2022073414A1