Intelligent patrol image semantic analysis and text generation method

Through deep learning and generative adversarial network (GAN) technology, sensitive areas in intelligent patrol images are identified and blurred, and image data is protected in combination with deep encryption technology, the risk of image data leakage is solved, and efficient privacy protection and secure transmission of image data is achieved.

CN119942545APending Publication Date: 2025-05-06CHINA SOUTHERN POWER GRID ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411898171.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the intelligent patrol image semantic analysis and text generation method, the image data contains a large amount of sensitive information, which poses the risk of malicious attack or unauthorized access in the process of transmission, storage and processing, resulting in the possibility of privacy violations and security accidents.

Method used

Deep learning and generative adversarial network (GAN) technology are used to identify sensitive areas in the image through convolutional neural network (CNN) and object detection algorithms and blur them. At the same time, deep encryption technology is used to ensure the privacy and security of image data. Combining AES-256 encryption, TLS/SSL protocol and blockchain technology, we ensure the security, privacy and immutability of generated text data during storage and transmission.

Benefits of technology

It effectively protects sensitive information in smart inspection images, reduces the risk of data leakage, ensures privacy protection during image acquisition and processing, and improves the credibility and security of the smart inspection system in terms of privacy protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942545A_ABST
    Figure CN119942545A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent patrol image semantic analysis and text generation method, and relates to the technical field of image semantic analysis and text generation, and the method comprises the following steps: collecting image data through an intelligent patrol device, and carrying out the preliminary processing of the collected image data, so as to improve the accuracy and stability of subsequent processing. According to the method, deep learning and generative adversarial network technologies are combined, and sensitive information in the intelligent patrol image is protected. A sensitive area is identified and fuzzified through a convolutional neural network, and data security is ensured in cooperation with a deep encryption technology. Image semantic analysis and text generation are combined with deep learning, the text quality is optimized, and the intelligent level is improved. In the data transmission and storage process, AES-256 encryption, a TLS / SSL protocol and a block chain technology are adopted to ensure the security, privacy and non-tampering of the text, and the reliability and compliance of the system are effectively guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image semantic analysis and text generation, and in particular to an intelligent patrol image semantic analysis and text generation method. Background Art

[0002] Intelligent patrol image semantic analysis and text generation refers to the use of computer vision and natural language processing technology to extract semantic information from images through in-depth analysis and processing of images collected during the patrol process, and convert this information into structured text descriptions. Specifically, semantic segmentation technology marks and classifies different areas in the image, target detection identifies and locates specific objects or events in the image, and image-to-text conversion converts the processed image information into natural language descriptions. This technology is widely used in the fields of intelligent monitoring, security patrol, etc., and can realize automated image analysis and report generation, greatly improving patrol efficiency and accuracy.

[0003] The prior art has the following deficiencies:

[0004] In the method of semantic parsing and text generation of intelligent patrol images, image data usually contains a large amount of sensitive information, such as personal faces, vehicle license plates, and building interiors. If the technology fails to effectively encrypt and anonymize data, there is a risk of malicious attacks or unauthorized third-party access during data transmission, storage, and processing. Once this sensitive information is leaked, it may lead to serious privacy violations and security incidents, especially in areas such as public security and financial monitoring, which may cause social panic, legal proceedings, and even economic losses. Therefore, how to protect data privacy and prevent information leakage while ensuring efficient image semantic parsing and text generation has become a major and urgent technical problem facing this technology.

[0005] The above information disclosed in this Background section is only for enhancement of understanding of the background of the present disclosure and therefore it may contain information that does not constitute the prior art that is already known to one of ordinary skill in the art. Summary of the invention

[0006] The purpose of the present invention is to provide a method for semantic parsing and text generation of intelligent patrol images, combining deep learning and generative adversarial network (GAN) technology to effectively protect sensitive information in intelligent patrol images. Sensitive areas in the image (such as faces, license plates, etc.) are identified through convolutional neural networks (CNN) and target detection algorithms, and GAN is used for fuzzy processing. At the same time, deep encryption technology is used to ensure the privacy and security of image data and reduce the risk of data leakage. Image semantic parsing and text generation are combined with deep learning technology to generate accurate and contextual natural language descriptions, optimize text quality through reinforcement learning, and improve the intelligence level of the system. During data transmission and storage, the present invention adopts AES-256 encryption, TLS / SSL protocol and blockchain technology to ensure the security, privacy and non-tamperability of the generated text, effectively ensuring the reliability and compliance of the intelligent patrol system to solve the problems in the above-mentioned background technology.

[0007] In order to achieve the above-mentioned purpose, the present invention provides the following technical solution: an intelligent patrol image semantic parsing and text generation method, which aims to solve the risk of sensitive information leakage in image data, comprising the following steps:

[0008] Collect image data through intelligent patrol equipment and perform preliminary processing on the collected image data, including removing noise and enhancing image quality, so as to improve the accuracy and stability of subsequent processing;

[0009] Before image data is stored and transmitted, encryption technology based on deep learning privacy protection algorithm is used to encrypt sensitive information, and facial blurring and license plate character masking are performed on identifiable people, vehicle license plates and other private information to prevent sensitive information leakage at the source;

[0010] Use convolutional neural network (CNN) and image segmentation algorithm to perform semantic analysis on the processed images, identify and annotate various objects and scene information in the images, and generate multi-level semantic labels, including the identification of people, objects, environment and abnormal events;

[0011] A text generation model based on a generative adversarial network (GAN) is used to combine the parsed image information with known semantic rules to generate a natural language description that is highly relevant to the image content and has logic; the text description includes but is not limited to a detailed description of the objects and events in the image, and has a sense of hierarchy to ensure the integrity and readability of the content;

[0012] To further ensure information security, after the text is generated, an advanced encryption algorithm is used to encrypt the generated text data to ensure that sensitive information will not be leaked during storage or transmission. The encrypted text data is only allowed to be decrypted and viewed by authorized users, thus ensuring data privacy and security.

[0013] Preferably, the image encryption and anonymization processing step includes the following sub-steps:

[0014] In the image acquisition stage, the real-time image is preliminarily processed, and all objects in the image are classified and areas that may contain sensitive information are identified through image feature extraction technology based on convolutional neural network (CNN);

[0015] For the identified sensitive information areas (such as faces, license plates, etc.), these areas are dynamically encrypted and blocked through the fuzzification technology based on the Generative Adversarial Network (GAN);

[0016] The entire image is encrypted using an encryption algorithm, converting the image data into encrypted form to ensure data security during transmission and storage, and prevent unauthorized users from accessing sensitive content.

[0017] Preferably, the image semantic analysis step includes the following:

[0018] When parsing the image, a target detection algorithm based on a deep convolutional neural network (CNN) is used to locate objects in the image and generate a bounding box for each target;

[0019] Use regional convolutional neural network (R-CNN) to more accurately classify objects within each bounding box, thereby improving the accuracy of object detection;

[0020] Based on image segmentation technology, the image is divided into different regions, and each region is assigned a corresponding semantic label to ensure that each part of the image can be reasonably semantically understood.

[0021] Preferably, the text generation step further includes the following:

[0022] Use a bidirectional long short-term memory network (Bi-LSTM) to combine multi-level semantic label information in the image to generate a natural language description that is highly relevant to the image content;

[0023] During the generation process, a strategy optimization mechanism based on reinforcement learning (RL) is used to conduct multiple rounds of evaluation and optimization on the generated text. The rationality of the generated text is fed back by the discriminant network to ensure the coherence and accuracy of the text.

[0024] Text generation not only includes information about objects, events, and environments in images, but can also generate dynamic inference information based on the context, thereby improving the richness and detail of text descriptions.

[0025] Preferably, the image encryption and anonymization processing step further includes:

[0026] When processing images, we use multi-level encryption technology, and use different encryption strengths for different areas according to the importance of the image content. For sensitive areas, such as faces, license plates, and important landmarks, we use higher levels of encryption, while for irrelevant areas, we use light encryption.

[0027] Use a privacy-preserving generative adversarial network (GAN) to generate image encryption templates to ensure that the privacy protection level during image encryption changes dynamically with different scenarios, thereby improving the adaptability and intelligence of the system;

[0028] During the encrypted image processing process, differential privacy technology is used to perturb the data, effectively increasing the privacy and anonymity of the data while ensuring that the data semantics are not affected.

[0029] Preferably, the text generation and multi-layer semantic association steps include the following features:

[0030] Based on the results of image semantic analysis, a graph neural network (GNN) is used to model various semantic objects in the image and their relationships, thereby ensuring that the generated text description has a clear structure and logic;

[0031] In the text generation process, deep neural networks (DNNs) are used to associate semantic information at different levels, thereby generating multi-dimensional natural language descriptions that can cover the detailed information in the image and fully demonstrate the relationship between objects;

[0032] In order to further improve the descriptive effect of the text, a multimodal learning algorithm is used to combine image and text information, and to improve the quality and accuracy of text generation through cross-modal association.

[0033] Preferably, the encrypted text transmission and storage step includes the following sub-steps:

[0034] After the text data is generated, it is first encrypted with high-strength AES-256 to ensure the integrity and confidentiality of the text content during storage and transmission;

[0035] For encrypted text data, ciphertext storage and transmission technology is used to prevent the data from being decrypted by unauthorized users even if it is intercepted during transmission, thus preventing the leakage of sensitive information;

[0036] During the transmission of encrypted text, blockchain technology is used for decentralized storage to ensure the immutability and credibility of text data, and provide a transparent audit mechanism to make the use and storage process of data traceable.

[0037] Preferably, the real-time monitoring and verification of the image and generated text includes the following specific steps:

[0038] In each step of image acquisition and text generation, intelligent monitoring algorithms are used to detect in real time whether there is a risk of data leakage or tampering, ensuring that the entire process meets privacy protection standards;

[0039] Every time a text is generated, a real-time detection module based on anomaly detection technology is used to monitor whether the generated text contains sensitive information or illegal behavior, so as to ensure that the generated text complies with privacy protection standards;

[0040] The image and text generation process is subject to strict security audits, and each step of the operation is recorded through real-time security logs for subsequent review and risk assessment.

[0041] In the above technical solution, the technical effects and advantages provided by the present invention are:

[0042] The present invention effectively protects sensitive information in intelligent patrol images by using encryption and anonymization technology that combines deep learning and generative adversarial networks (GAN). For private areas such as faces, license plates, and building interiors in the image, the system accurately identifies and locates these sensitive areas through convolutional neural networks (CNN) and target detection algorithms, then uses GAN to blur them, and further encrypts the entire image through deep encryption technology. Through this comprehensive processing method, even if the image data is intercepted, unauthorized personnel cannot restore sensitive information, avoiding the risk of data leakage and ensuring privacy protection during image acquisition and processing. Especially in high-risk areas such as public security and financial monitoring, this technology can effectively reduce the possibility of data leakage and improve the credibility and security of the intelligent patrol system in terms of privacy protection.

[0043] The present invention adopts a multi-level fusion image semantic parsing and text generation method, which significantly improves the accuracy and fluency of text generation. Through the combination of deep convolutional neural network (CNN) and regional convolutional neural network (R-CNN), the system can efficiently identify a variety of objects and scenes in the image, and provide accurate positioning and classification for each object. Subsequently, based on the generative adversarial network (GAN) and the bidirectional long short-term memory network (Bi-LSTM), the text description generated by the system can not only accurately reflect the objects in the image, but also generate natural language text that conforms to the context according to the context in the image. The generation process is optimized by reinforcement learning (RL), and the quality of the text description is further improved, ensuring that the generated text performs well in terms of logic, coherence and semantic integrity. This technology can be applied in multiple fields, such as intelligent inspection, security monitoring and other scenes, automatically generating high-quality image descriptions and reports, and significantly improving the intelligence level and work efficiency of the system.

[0044] The present invention effectively ensures the security, privacy and non-tamperability of the generated text during storage and transmission by combining advanced encryption technology and blockchain technology. The generated text is encrypted and stored using the AES-256 encryption algorithm and securely transmitted using the TLS / SSL protocol, preventing data leakage and tampering during transmission and storage. At the same time, the text data is decentralized and stored using blockchain to ensure that the text will not be illegally modified or deleted during storage, and each piece of text data can trace its generation and access history. This technology provides a complete data security mechanism and transparent auditing function, ensuring the integrity and legitimacy of text data in sensitive areas such as intelligent patrols and security monitoring. This data security measure greatly enhances the credibility of the system and ensures that the text generation and transmission process meets privacy protection and compliance requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0046] Figure 1 The present invention is a method flow chart of the intelligent patrol image semantic analysis and text generation method. DETAILED DESCRIPTION

[0047] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these example embodiments are provided so that the description of the present disclosure will be more comprehensive and complete, and the concept of the example embodiments will be fully conveyed to those skilled in the art.

[0048] The present invention provides Figure 1 The method for semantic parsing and text generation of intelligent patrol images is designed to address the risk of sensitive information leakage in image data, and includes the following steps:

[0049] Collect image data through intelligent patrol equipment and perform preliminary processing on the collected image data, including removing noise and enhancing image quality, so as to improve the accuracy and stability of subsequent processing;

[0050] Before image data is stored and transmitted, encryption technology based on deep learning privacy protection algorithm is used to encrypt sensitive information, and facial blurring and license plate character masking are performed on identifiable people, vehicle license plates and other private information to prevent sensitive information leakage at the source;

[0051] Use convolutional neural network (CNN) and image segmentation algorithm to perform semantic analysis on the processed images, identify and annotate various objects and scene information in the images, and generate multi-level semantic labels, including the identification of people, objects, environment and abnormal events;

[0052] A text generation model based on a generative adversarial network (GAN) is used to combine the parsed image information with known semantic rules to generate a natural language description that is highly relevant to the image content and has logic; the text description includes but is not limited to a detailed description of the objects and events in the image, and has a sense of hierarchy to ensure the integrity and readability of the content;

[0053] To further ensure information security, after the text is generated, an advanced encryption algorithm is used to encrypt the generated text data to ensure that sensitive information will not be leaked during storage or transmission. The encrypted text data is only allowed to be decrypted and viewed by authorized users, thus ensuring data privacy and security.

[0054] Specific implementation method 1: Image encryption and anonymization processing based on deep learning;

[0055] In this embodiment, first, intelligent patrol equipment (such as drones, fixed cameras or handheld devices) will collect images in a predetermined area. These devices have high real-time image processing capabilities and can quickly collect high-quality image data under complex environmental conditions. The collected images are first preprocessed, including removing noise, adjusting contrast, enhancing image clarity, etc. These processing steps are intended to improve the accuracy of subsequent analysis and processing. The preprocessed images provide a good basis for subsequent sensitive information detection and encryption.

[0056] Next, sensitive areas in the image are automatically identified and marked. A convolutional neural network (CNN) is used to perform a preliminary analysis of the image. CNN can effectively extract features from the image and classify and locate objects in the image through a trained model. For example, in an image, if there are people, license plates, or sensitive information inside a building, CNN will locate these areas through object detection algorithms (such as YOLO or Faster R-CNN) and identify faces, license plates, and other sensitive areas.

[0057] The encryption and anonymization of the identified sensitive areas are the core of this implementation. First, the sensitive areas are dynamically processed using a blurring technique based on a generative adversarial network (GAN). The GAN consists of a generator and a discriminator, where the generator is responsible for generating a blur effect for processing the image, while the discriminator determines whether the generated blur effect is sufficient to conceal sensitive information. Through continuous training and optimization, the generator can generate an effective blurring effect to ensure that information such as faces and license plates cannot be restored or inferred.

[0058] After blurring, in order to further enhance privacy protection, the entire image data will be encrypted using a deep learning encryption algorithm. Specifically, the encryption method used in this embodiment is an encryption network based on a convolutional neural network (CNN), which can dynamically adjust the encryption strength according to the complexity of the image content, using a lower encryption strength for simple image areas and a stronger encryption for complex sensitive areas. The encrypted image cannot be accessed or decrypted by unauthorized users. Even if a hacker intercepts the encrypted data, the original image and the sensitive information it contains cannot be restored.

[0059] In addition, during the image transmission process, all image data will use secure transmission protocols, such as TLS / SSL and other encrypted communication protocols, to ensure the confidentiality and integrity of the data during transmission. During the storage process, the image data will be stored in an encrypted form on the server or cloud platform, and only authorized users can access the encrypted image data. This encrypted storage and transmission solution effectively prevents the leakage and tampering of image data.

[0060] Finally, the system will provide real-time monitoring and alert functions to ensure that if data leakage or illegal access risks are found at any stage, intervention and repair can be carried out immediately. The monitoring system will record data storage and transmission operations in real time and generate logs for subsequent audits and risk assessments.

[0061] In summary, this implementation can effectively protect sensitive information in intelligent patrol images by combining deep learning, generative adversarial networks (GANs) and encryption technology, ensuring that no privacy information is leaked during data transmission, storage and processing. Whether it is image acquisition, encryption processing, or data transmission and storage, the integrity and security of information can be guaranteed, providing reliable privacy protection for the semantic analysis of intelligent patrol images.

[0062] Specific implementation method 2: Multi-level fusion of image semantic analysis and text generation;

[0063] This implementation method focuses on the multi-level integration of image semantic analysis and text generation, uses deep learning technology to perform semantic analysis on images, and generates high-quality natural language descriptions. First, the image data collected by the intelligent patrol equipment will be preprocessed to remove noise and enhance image clarity. The preprocessed image will be sent to a deep convolutional neural network (CNN) for preliminary analysis. CNN can extract feature information from the image and identify objects and scenes therein. The image is analyzed through target detection algorithms (such as YOLO and Faster R-CNN). CNN can identify a variety of objects such as people, vehicles, buildings, plants, etc. that appear in the image.

[0064] After the recognition and location of each object in the image are completed, the system will use the regional convolutional neural network (R-CNN) to classify each object more accurately and generate bounding boxes for each object. The objects in each bounding box will be further classified, such as "pedestrian", "car", "doorway" or "room", etc. These labels will become the basis for subsequent text generation. The semantic information of the image is further processed through image segmentation technology, which divides the image into multiple semantic regions, each of which represents a specific object or environment in the image.

[0065] Next, the semantic information of the image is input into the text generation module. In this embodiment, the text generation uses a method based on the Generative Adversarial Network (GAN). The Generative Adversarial Network consists of two main parts: the generator and the discriminator. The generator generates a natural language description based on the semantic labels of the image, describing each object in the image and their relationship with each other; the discriminator evaluates the generated text to determine whether the text is true and reasonable, and continuously optimizes the output of the generator.

[0066] In addition to using a generative adversarial network (GAN), this implementation also combines a bidirectional long short-term memory network (Bi-LSTM) to handle the text generation process. Bi-LSTM can capture the contextual relationships in the image through a bidirectional information flow and generate a more natural and contextual language description. During the generation process, Bi-LSTM will generate coherent and logical text based on the contextual relationships in the image, which can not only accurately describe the location and type of the object, but also better describe the activities, events and relationships between objects in the image.

[0067] In order to improve the accuracy and richness of the text, this implementation also introduces a reinforcement learning (RL) mechanism. After each text generation, the discriminator will evaluate the text, identify the deficiencies in the text, and adjust the generator through a feedback mechanism. The reinforcement learning algorithm can dynamically adjust the model according to the quality of the generated text, continuously improving the accuracy, fluency, and information richness of the generated text.

[0068] Finally, the generated text description will be encrypted and transmitted to a secure storage system. In order to ensure the privacy and security of the text data, the generated text will be encrypted by an advanced encryption algorithm and stored in an encrypted form in a secure server to ensure that unauthorized users cannot access the generated text content.

[0069] This implementation effectively improves the quality and accuracy of image semantic analysis and text generation by combining deep convolutional neural networks (CNN), generative adversarial networks (GAN), bidirectional long short-term memory networks (Bi-LSTM) and reinforcement learning (RL) algorithms. The fusion of multi-level semantic analysis of images and text generation ensures that the generated text not only contains information about objects and events in the image, but also can reasonably infer the relationship between objects, providing strong technical support for automatic report generation during intelligent patrols.

[0070] Specific implementation method three: secure transmission and storage of encrypted text;

[0071] In this embodiment, the focus is on solving the problem of secure transmission and storage of generated text. First, after the image semantic analysis and text generation are completed, the system will perform advanced encryption on the generated text data to ensure that the text content will not be leaked. The text encryption adopts the AES-256 encryption algorithm, which is a recognized high-intensity encryption algorithm that can effectively prevent unauthorized access and cracking. After the text data is encrypted, it is stored in an encrypted database and transmitted through a secure communication protocol. In order to ensure the security during the data transmission process, all generated text data is encrypted and transmitted via the TLS / SSL protocol, which can ensure the integrity and privacy of the data when it is transmitted on the Internet.

[0072] The encrypted text will be stored in the cloud platform or local server. During the storage process, this implementation adopts a decentralized storage solution and combines blockchain technology to store and manage text data. The decentralized nature of blockchain can ensure the immutability and transparency of data. Each generated text will be timestamped and recorded in the blockchain to ensure that the text content will not be illegally modified or deleted during the storage process. At the same time, blockchain can also provide an audit tracking function for each piece of text data, which is convenient for security auditing and backtracking when needed.

[0073] In order to ensure that only authorized users can access the stored text data, this implementation also sets up a strict identity authentication mechanism. When users access text data, they need to provide valid identity authentication credentials, such as user name, password, fingerprint, facial recognition and other multiple authentication methods. These authentication methods ensure that only authorized personnel can decrypt and access text data, and unauthorized personnel will not be able to obtain any information.

[0074] During the entire text data storage and transmission process, this implementation provides a complete security monitoring mechanism that can monitor all operations in real time. All storage, transmission, and access operations of text data will be recorded in detail and logs will be generated to ensure that the operations are traceable. This log information can be used for subsequent security audits, risk assessments, and compliance checks.

[0075] This implementation ensures the security, privacy, and integrity of the generated text during storage, transmission, and access by combining the AES-256 encryption algorithm, TLS / SSL protocol, blockchain technology, and multi-factor authentication.

[0076] The present invention effectively protects sensitive information in intelligent patrol images by using encryption and anonymization technology that combines deep learning and generative adversarial networks (GAN). For private areas such as faces, license plates, and building interiors in the image, the system accurately identifies and locates these sensitive areas through convolutional neural networks (CNN) and target detection algorithms, then uses GAN to blur them, and further encrypts the entire image through deep encryption technology. Through this comprehensive processing method, even if the image data is intercepted, unauthorized personnel cannot restore sensitive information, avoiding the risk of data leakage and ensuring privacy protection during image acquisition and processing. Especially in high-risk areas such as public security and financial monitoring, this technology can effectively reduce the possibility of data leakage and improve the credibility and security of the intelligent patrol system in terms of privacy protection.

[0077] The present invention adopts a multi-level fusion image semantic parsing and text generation method, which significantly improves the accuracy and fluency of text generation. Through the combination of deep convolutional neural network (CNN) and regional convolutional neural network (R-CNN), the system can efficiently identify a variety of objects and scenes in the image, and provide accurate positioning and classification for each object. Subsequently, based on the generative adversarial network (GAN) and the bidirectional long short-term memory network (Bi-LSTM), the text description generated by the system can not only accurately reflect the objects in the image, but also generate natural language text that conforms to the context according to the context in the image. The generation process is optimized by reinforcement learning (RL), and the quality of the text description is further improved, ensuring that the generated text performs well in terms of logic, coherence and semantic integrity. This technology can be applied in multiple fields, such as intelligent inspection, security monitoring and other scenes, automatically generating high-quality image descriptions and reports, and significantly improving the intelligence level and work efficiency of the system.

[0078] The present invention effectively ensures the security, privacy and non-tamperability of the generated text during storage and transmission by combining advanced encryption technology and blockchain technology. The generated text is encrypted and stored using the AES-256 encryption algorithm and securely transmitted using the TLS / SSL protocol, preventing data leakage and tampering during transmission and storage. At the same time, the text data is decentralized and stored using blockchain to ensure that the text will not be illegally modified or deleted during storage, and each piece of text data can trace its generation and access history. This technology provides a complete data security mechanism and transparent auditing function, ensuring the integrity and legitimacy of text data in sensitive areas such as intelligent patrols and security monitoring. This data security measure greatly enhances the credibility of the system and ensures that the text generation and transmission process meets privacy protection and compliance requirements.

[0079] The above description is only by way of illustration of certain exemplary embodiments of the present invention. It is undoubted that those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the above drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.

[0080] It should be noted that, in this article, if there are relational terms such as first and second, etc., they are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0081] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0082] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0083] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0084] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0085] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0086] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0087] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage media include: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks or optical disks.

[0088] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. Intelligent patrol image semantic analysis and text generation method, characterized in that: The following steps are involved: Collect image data through intelligent patrol equipment and perform preliminary processing on the collected image data to improve the accuracy and stability of subsequent processing; Before image data is stored and transmitted, encryption technology based on deep learning privacy protection algorithm is used to encrypt sensitive information, and the faces of identifiable people and vehicle license plates are blurred and the license plate characters are blocked to prevent sensitive information leakage from the source; Use convolutional neural networks and image segmentation algorithms to perform semantic analysis on the processed images, identify and annotate various objects and scene information in the images, and generate multi-level semantic labels, including the identification of people, objects, environments, and abnormal events; A text generation model based on generative adversarial networks is used to combine the parsed image information with known semantic rules to generate a natural language description that is highly relevant to the image content and has logic. In the process of text generation, advanced encryption algorithms are used to encrypt the generated text data to ensure that sensitive information will not be leaked during storage or transmission. The encrypted text data is only allowed to be decrypted and viewed by authorized users to ensure data privacy and security.

2. The method for semantic analysis and text generation of intelligent patrol images according to claim 1 is characterized in that: The image encryption and anonymization process includes the following sub-steps: In the image acquisition stage, the real-time image is preliminarily processed, and all objects in the image are classified and areas containing sensitive information are identified through image feature extraction technology based on convolutional neural networks; For the identified sensitive information areas, the fuzzification technology based on the generative adversarial network is used to dynamically encrypt and mask the sensitive information areas; The entire image is encrypted using an encryption algorithm, converting the image data into encrypted form to ensure data security during transmission and storage, and prevent unauthorized users from accessing sensitive content.

3. The method for semantic analysis and text generation of intelligent patrol images according to claim 1 is characterized in that: The steps of image semantic parsing include the following: When parsing the image, a target detection algorithm based on a deep convolutional neural network is used to locate objects in the image and generate a bounding box for each target; Use regional convolutional neural networks to more accurately classify objects within each bounding box, thereby improving the accuracy of object detection; Based on image segmentation technology, the image is divided into different regions, and each region is assigned a corresponding semantic label to ensure that each part of the image can be reasonably semantically understood.

4. The method for semantic analysis and text generation of intelligent patrol images according to claim 1 is characterized in that: The text generation step further includes the following: Use a bidirectional long short-term memory network to combine multi-level semantic label information in the image to generate a natural language description that is highly relevant to the image content; During the generation process, a strategy optimization mechanism based on reinforcement learning is used to conduct multiple rounds of evaluation and optimization on the generated text. The rationality of the generated text is fed back by the discriminant network to ensure the coherence and accuracy of the text. Text generation not only includes information about objects, events, and the environment in the image, but can also generate dynamic inference information based on the context, thereby improving the richness and detail of the text description.

5. The method for semantic analysis and text generation of intelligent patrol images according to claim 1 is characterized in that: The image encryption and anonymization processing steps further include: When processing images, multi-level encryption technology is used. Different encryption strengths are used for different areas according to the importance of the image content. A higher level of encryption is used for sensitive areas, while light encryption is used for irrelevant areas. Use a privacy-preserving generative adversarial network to generate image encryption templates, ensuring that the privacy protection level during image encryption changes dynamically with different scenarios to improve the adaptability and intelligence of the system; During the encrypted image processing, differential privacy technology is used to perturb the data, thereby increasing the privacy and anonymity of the data while ensuring that the data semantics are not affected.

6. The method for semantic analysis and text generation of intelligent patrol images according to claim 1 is characterized in that: The text generation and multi-layer semantic association steps include the following features: Based on the results of image semantic analysis, a graph neural network is used to model various semantic objects in the image and their relationships, ensuring that the generated text description has a clear structure and logic. In the text generation process, deep neural networks are used to associate semantic information at different levels to generate multi-dimensional natural language descriptions that cover the detailed information in the image and fully demonstrate the relationship between objects; In order to further improve the descriptive effect of the text, a multimodal learning algorithm is used to combine image and text information, and to improve the quality and accuracy of text generation through cross-modal association.

7. The method for semantic analysis and text generation of intelligent patrol images according to claim 1 is characterized in that: The encrypted text transmission and storage step includes the following sub-steps: After the text data is generated, it is first encrypted with high-strength AES-256 to ensure the integrity and confidentiality of the text content during storage and transmission; For encrypted text data, ciphertext storage and transmission technology is used to prevent the data from being decrypted by unauthorized users even if it is intercepted during transmission, thus preventing the leakage of sensitive information; During the transmission of encrypted text, blockchain technology is used for decentralized storage to ensure the immutability and credibility of text data, and provide a transparent audit mechanism to make the use and storage process of data traceable.

8. The method for semantic analysis and text generation of intelligent patrol images according to claim 1 is characterized in that: Real-time monitoring and verification of images and generated text. The specific steps include: In each step of image acquisition and text generation, intelligent monitoring algorithms are used to detect in real time whether there is a risk of data leakage or tampering, ensuring that the entire process meets privacy protection standards; Every time a text is generated, a real-time detection module based on anomaly detection technology is used to monitor whether the generated text contains sensitive information or illegal behavior, so as to ensure that the generated text complies with privacy protection standards; The image and text generation process is subject to strict security audits, and each step of the operation is recorded through real-time security logs for subsequent review and risk assessment.

Citation Information

Cited By

  • Picture information broadcasting method and device, electronic equipment and storage medium

    CN121418608A