Intelligent networked automobile spatio-temporal data processing method, device and equipment and storage medium

By extracting frames and identifying data packets of intelligent connected vehicles, the problem of desensitizing geographically sensitive information in spatiotemporal data acquired over a large area and for a long time is solved, achieving efficient and secure data processing.

CN120635871APending Publication Date: 2025-09-12BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202411930161.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

In intelligent connected vehicles, how to effectively process spatiotemporal data acquired over a large area, for a long time, and at multiple levels to maintain the security of surveying and mapping geographic information, especially how to desensitize geographic sensitive information.

Method used

By extracting frames from data packets collected by intelligent connected vehicles, we obtain multiple frames of images with an overlap of more than 50% between two adjacent frames, identify the geographically sensitive information in these images, and mark the data packets based on the identification results.

Benefits of technology

It reduces the computational complexity of sensitive information identification, improves the tolerance of identification errors, ensures the security of geographic information, reduces the possibility of omission of sensitive information, and realizes the effective processing of geographic sensitive information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635871A_ABST
    Figure CN120635871A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent networked automobile spatio-temporal data processing method, device and equipment and a storage medium, and relates to the technical field of computers, in particular to the technical field of intelligent networked automobiles, geographic information surveying and mapping and the like. According to the specific implementation scheme, frame extraction is carried out on a data packet collected by the intelligent networked automobile to obtain multiple frames of images, and the overlapping degree of every two adjacent frames of images in the multiple frames of images is larger than or equal to 50%; based on the multiple frames of images, geographic sensitive information in the multiple frames of images is identified; and marking the data packet according to an identification result. According to the invention, the sensitive information in the spatio-temporal data of the intelligent networked automobile can be identified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to technical fields such as intelligent connected vehicles and surveying and mapping geographic information. Background Art

[0002] In the field of intelligent connected vehicle technology, intelligent connected vehicles acquire and aggregate spatiotemporal data over a large range, for a long time, and at multiple levels. How to desensitize the spatiotemporal data generated by intelligent connected vehicles to maintain the security of surveying and mapping geographic information is a technical problem that needs to be solved. Summary of the Invention

[0003] The present disclosure provides a method, apparatus, device, and storage medium for processing spatiotemporal data of an intelligent connected vehicle.

[0004] According to one aspect of the present disclosure, a method for processing spatiotemporal data of an intelligent connected vehicle is provided, comprising:

[0005] Extracting frames from data packets collected by the intelligent connected vehicle to obtain a multi-frame image, wherein the overlap between two adjacent frames of the multi-frame image is greater than or equal to 50%;

[0006] Based on the multiple frames of images, identifying geographically sensitive information in the multiple frames of images;

[0007] The data packet is marked according to the identification result.

[0008] According to another aspect of the present disclosure, a device for processing spatiotemporal data of an intelligent connected vehicle is provided, comprising:

[0009] A frame extraction module is used to extract frames from data packets collected by the intelligent connected vehicle to obtain multiple frames of images, wherein the overlap between two adjacent frames of images in the multiple frames is greater than or equal to 50%;

[0010] An identification module, configured to identify geographically sensitive information in the multiple image frames based on the multiple image frames;

[0011] The marking module is used to mark the data packet according to the recognition result.

[0012] According to another aspect of the present disclosure, there is provided an electronic device, comprising:

[0013] at least one processor; and

[0014] a memory communicatively connected to the at least one processor; wherein,

[0015] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform any method in the embodiments of the present disclosure.

[0016] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any method according to the embodiments of the present disclosure.

[0017] According to another aspect of the present disclosure, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the computer program implements any one of the methods according to the embodiments of the present disclosure.

[0018] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0020] Figure 1 This is a flowchart of a method for processing spatiotemporal data of an intelligent connected vehicle according to an embodiment of the present disclosure;

[0021] Figure 2 is a functional diagram of each processing module according to an embodiment of the present disclosure;

[0022] Figure 3 It is a schematic diagram of geographically sensitive information image features Figure 1 ;

[0023] Figure 4 It is a schematic diagram of geographically sensitive information image features Figure 2 ;

[0024] Figure 5 is a structural diagram of a spatiotemporal data processing device 500 for an intelligent connected vehicle according to an embodiment of the present disclosure;

[0025] Figure 6 A schematic block diagram of an example electronic device 600 is shown, which may be used to implement embodiments of the present disclosure. DETAILED DESCRIPTION

[0026] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0027] The “and / or” in the embodiments of the present disclosure indicates that there may be three relationships. For example, A and / or B may indicate three situations: A exists alone, A and B exist at the same time, and B exists alone. The term “at least one” herein indicates any combination of at least two of any one or more of a plurality of. For example, at least one of A, B, and C may indicate any one or more elements selected from the set consisting of A, B, and C. The terms “first” and “second” herein refer to and distinguish between multiple similar technical terms, and do not mean to limit the order or to limit the meaning to only two. For example, the first feature and the second feature refer to two categories / two features. The first feature may be one or more, and the second feature may also be one or more.

[0028] In the field of intelligent connected vehicle technology, intelligent connected vehicles acquire and aggregate spatiotemporal data over a large area, over a long period of time, and at multiple levels. This spatiotemporal data is stored in the form of data packets (or task packets). Different automakers have different sensor configurations on the vehicle side of intelligent connected vehicles, and the image data collected by the vehicle side has different specifications. For example, the data collected by the vehicle side comes in various formats, such as YUV, RAW, Automotive Data and Time-Triggered Framework (ADTFADT), and Image data (IMG). The disclosed embodiments can parse the data collected by the vehicle side (video data YUV, RAW, and specific formats: ADTFADT, ROSBAG) to obtain image information. For example, video data (such as Moving Picture Experts Group-4 (MP4), Audio Video Interleaved (AVI), and High Efficiency Video Coding (H265)) can be obtained. Currently, intelligent connected vehicles are generally equipped with more than six cameras and use a trigger-based acquisition strategy. The time range of each data packet is generally 10s to 60s, and the video frame rate of the data packet is generally in the range of 10fps to 30fps (frames per second).

[0029] The disclosed embodiment proposes a method for processing spatiotemporal data of an intelligent connected vehicle, which can mark sensitive information in a data packet of an intelligent connected vehicle so as to process the sensitive information in the data packet and prevent the sensitive information from being disclosed.

[0030] Figure 1 This is a flowchart of a method for processing spatiotemporal data of an intelligent connected vehicle according to an embodiment of the present disclosure, including:

[0031] S110, extracting frames from data packets collected by the intelligent connected vehicle to obtain multiple frames of images, wherein an overlap between two adjacent frames of images in the multiple frames of images is greater than or equal to 50%;

[0032] S120: Identify geographically sensitive information in the multiple image frames based on the multiple image frames;

[0033] S130: Mark the data packet according to the identification result.

[0034] The disclosed embodiment extracts frames from data packets collected by intelligent connected vehicles, and then identifies geographically sensitive information in the extracted images, thereby enabling sensitive data tagging of data packets. Since the number of images after extraction is less than the number of images in the entire data packet, the amount of computation required to identify sensitive information can be reduced. Since the degree of overlap between two adjacent frames of images after extraction is greater than or equal to 50%, it is ensured that all contents in the extracted images can appear in the two adjacent frames; that is, for any image element, sensitive information identification is performed twice (identification is performed on the two frames of images containing the image element separately), thereby improving the tolerance for identification errors.

[0035] The disclosed embodiments implement desensitizing processing of spatiotemporal data acquired and aggregated over a large area, for a long time, and at multiple levels for intelligent connected vehicles, thereby maintaining the security of surveying and mapping geographic information and promoting the development of intelligent connected vehicles.

[0036] In some embodiments, the geographically sensitive information includes information about at least one of the following units and facilities:

[0037] Dedicated railways and train lines within stations, railway marshalling yards, and dedicated roads.

[0038] Different automakers use different on-board sensor configurations in their smart connected vehicles, and the image data they collect varies in format. Before desensitizing, the disclosed embodiments can parse the collected data (such as YUV, RAW, adtfadt, and Img) to obtain corresponding data packets. These data packets can be in formats such as MP4, AVI, and H265. Currently, mainstream smart connected vehicles are typically equipped with six or more cameras and employ a trigger-based acquisition strategy. Each data packet typically ranges from 10s to 60s, and the video frame rate typically ranges from 10fps to 30fps (frames per second). Given this situation, the disclosed embodiments utilize step S110 to extract frames from the data packets collected by the smart connected vehicle and identify geo-sensitive information from the resulting multiple frames. Because adjacent images after extraction maintain a certain degree of overlap, this reduces computational cost and minimizes the likelihood of missing sensitive information, thus achieving a balance between cost and performance. Extracting the data packets yields multiple single-frame images and their corresponding timestamps.

[0039] In some embodiments, the frame extraction method is related to the driving speed of the intelligent connected vehicle and the video frame rate of the data packet. For example, in urban road scenarios, the speed of the car is generally 10m / s to 20m / s. When the vehicle camera captures the video frame rate at 20fps, the distance between two adjacent frames in the data packet is 0.5m to 1m; taking the effective coverage distance of a single image as 10 meters as an example, the overlap between two adjacent frames in the data packet is 90% to 95%. The data packet is frame extracted, and the overlap between two adjacent frames after frame extraction is greater than or equal to 50%. Therefore, the frame extraction method is related to the driving speed of the intelligent connected vehicle and the video frame rate of the data packet. Its essence is that the frame extraction method is related to the overlap between two adjacent frames in the data packet. For example, one frame is extracted every N frames; the higher the overlap between two adjacent frames in the data packet, the larger the value of N can be; the lower the overlap between two adjacent frames in the data packet, the smaller the value of N can be; to ensure that the overlap between two adjacent frames after frame extraction is greater than or equal to 50%, thereby reducing the amount of calculation and the probability of failure to recall sensitive information due to misidentification.

[0040] In one example, N is 2, 3, 4, or 5. In the aforementioned case, taking N=2 as an example, the overlap rate of two adjacent frames after frame extraction is 80% to 90%; taking N=5 as an example, the overlap rate of two adjacent frames after frame extraction is 50% to 75%.

[0041] The embodiment of the present disclosure defines the image features of geographically sensitive information based on the target data that must not be stored as stipulated in relevant regulations, and adopts an image recognition method to perform corresponding recognition of the image features of geographically sensitive information, and returns information such as the recognition category, recognition elements, element coordinates, and confidence level.

[0042] Among them, geographically sensitive information image features may include sensitive unit image features and sensitive facility image features.

[0043] The embodiment of the present disclosure can perform at least one of image recognition and text recognition on the multiple frames of images obtained by frame extraction to identify at least one of text-sensitive elements and image-sensitive elements related to geographic sensitive information.

[0044] In some embodiments, an image recognition method (such as an image recognition model) can be used to identify the image features of geographically sensitive information and return information such as the identification category, identification elements, element coordinates, and confidence level. A text recognition method, such as an optical character recognition (OCR) method, can be used to identify geographically sensitive information and text in an image to determine the terms contained in the image, the image coordinates of the terms, and the confidence level of the terms. For example, the term "First Hospital" in the image, the coordinates of the image of the term in the image, and the confidence level of the term can be identified. Based on the identified terms and a pre-set matching vocabulary, the sensitive words present in the image can be determined.

[0045] For the unit and facility attribute information identified by the image recognition method, and the sensitive words identified by the text recognition method and the pre-set matching word library, the geographically sensitive elements (including text-sensitive elements and image-sensitive elements) contained in the image can be determined based on the pre-set recall strategy. Afterwards, the determined geographically sensitive elements can be manually verified, and then the geographically sensitive elements determined by the manual verification can be processed. Specifically, the method of identifying geographically sensitive information includes: performing at least one of image recognition and text recognition on the multiple frames of images obtained after frame extraction; based on the results of at least one of the image recognition and text recognition, and the recall strategy for recalling geographically sensitive information, identifying geographically sensitive information in the multiple frames of images.

[0046] In some embodiments, for image recognition and text recognition, when the confidence level of the recognition result is greater than or equal to a preset threshold, the geographically sensitive information in the data packet is identified based on the recognition result and the recall strategy for recalling geographically sensitive information, that is, whether the recognition result belongs to geographically sensitive information is determined. In the case where the confidence level of the recognition result is less than the preset threshold, the recognition result is considered unreliable, and the subsequent determination of whether the recognition result belongs to sensitive information is no longer made. The embodiment of the present disclosure assumes that the recognition result does not belong to sensitive information. In one example, the preset threshold is 90%. For example, when the OCR recognition method is used to identify the existence of the term "First Hospital" in the image, and the confidence level of the recognition result is 80%, which is less than the preset threshold, the term will not be identified as sensitive information.

[0047] In some embodiments, the recall strategy includes at least one of the following:

[0048] (1) When the degree of match between the result of text recognition and the detection vocabulary is greater than or equal to a first threshold, and the degree of match between the result of text recognition and the exemption vocabulary is less than a second threshold, the object of text recognition is determined to be a text sensitive element;

[0049] (2) When the result of image recognition corresponds to geographically sensitive information, the object of image recognition is determined as an image sensitive element;

[0050] (3) When the recognition frame of the image recognition is smaller than or equal to the first size, or the recognition frame of the text recognition is smaller than or equal to the first size, the object of the image recognition or text recognition is determined as a non-sensitive element.

[0051] For the recall strategy in item (1) above, a detection word library and an exemption word library can be pre-set. The detection word library is used to identify sensitive content in images. The exemption word library is mainly used to exclude those words that appear in the image but are reasonable expressions in specific scenarios and the image should not be marked as sensitive.

[0052] In some embodiments, the first threshold is greater than or equal to 60%. For example, the first threshold is 75%. The value of the first threshold is related to the workload of the verification personnel in the subsequent manual verification process, as well as security. If the first threshold is set to 50%, when the result obtained by text recognition is half the same as the detection word in the detection word library, the result obtained by the text recognition will be determined as a text-sensitive element; in this case, more misjudgments will occur, resulting in a larger workload for subsequent verification personnel. For example, if the result obtained by text recognition is a word containing 2 characters or a word containing 4 characters, then when 1 character or 2 characters in the result appear in the search term, the result will be determined as a text-sensitive element.

[0053] Therefore, in order to avoid excessive workload for subsequent verification personnel and to ensure a certain recall rate to meet the security requirements for sensitive information identification, this solution sets the first threshold to be greater than or equal to 60%.

[0054] In some embodiments, the second threshold is 100%. Setting the second threshold can determine entries that completely match words in the exemption word library as non-sensitive words, thereby avoiding misidentification of reasonable expressions in some specific scenarios as sensitive words.

[0055] For example, if an OCR method is used to identify a term contained in an image, and the confidence level of the recognition result is 95%, which is greater than a preset threshold (e.g., 90%), the term is matched using the detection vocabulary and the exemption vocabulary. If the match between the term and the terms in the detection vocabulary is 100%, which is greater than a preset first threshold, and the match between the term and the relevant terms in the exemption vocabulary is 0%, which is less than a second threshold, the text recognition object is determined to be a text-sensitive element.

[0056] For example, if an OCR method is used to identify a term contained in an image, and the confidence level of the recognition result is 95%, which is greater than a preset threshold (e.g., 90%), the term is matched using the detection vocabulary and the exemption vocabulary. If the match between the term and the terms in the detection vocabulary is 80% (4 out of 5 characters match), which is greater than a preset first threshold, and the match between the term and the relevant terms in the exemption vocabulary is 100%, which is equal to a second threshold, then the text recognition object is determined not to be a sensitive text element.

[0057] For another example, using the OCR method to identify the words contained in the image, and the confidence level of the recognition result is 80%, which is less than a preset threshold (such as 90%), it is determined that the object of the text recognition is not a text-sensitive element.

[0058] By using the detection and exemption word library to detect sensitive words, we can not only identify words that are not allowed to appear in the public spatiotemporal data of intelligent connected vehicles, but also exclude those reasonable expressions that appear in the spatiotemporal data but belong to specific scenarios, thereby improving the accuracy of sensitive word identification.

[0059] In addition, this recall strategy only performs sensitive word recognition on entries whose text recognition confidence is greater than or equal to a preset threshold, thus avoiding the problem of sensitive word recognition errors caused by low text recognition confidence.

[0060] For the recall strategy in item (2) above, when the confidence level of the image recognition result is greater than or equal to the preset threshold, if the image recognition result corresponds to geographic sensitive information, the object of the image recognition is determined to be an image sensitive element.

[0061] For example, for recognition results such as a camouflaged house, the preset threshold can be set to 80%, that is, if the confidence of the recognition result is greater than or equal to 80% and the recognition result is a camouflaged house, the recognition result is determined to be an image sensitive element.

[0062] This recall strategy only recognizes image sensitive elements for objects whose image recognition confidence is greater than or equal to a preset threshold, thus avoiding the problem of incorrect recognition of image sensitive elements due to low image recognition confidence.

[0063] The recall strategy of item (3) above belongs to the minimum detected image element strategy, that is, when the detected feature element is too small, it means that the feature of the identification element is not clear and it may not be marked as a geographically sensitive element. In other words, the object of image recognition or text recognition can be determined as a non-geographically sensitive element. Specifically, when the object of image recognition or text recognition is a geographically sensitive element with a volume greater than a preset threshold, the first size is set to 20 pixels wide and 60 pixels high; or 60 pixels wide and 20 pixels high. That is, if the recognition box of a geographically sensitive element (such as a cooling tower) with a volume greater than the preset threshold is less than or equal to 20 pixels * 60 pixels, or less than or equal to 60 pixels * 20 pixels, the object of image recognition will not be marked as a geographically sensitive element.

[0064] For another example, if the image recognition object is a geo-sensitive element with a volume less than or equal to a preset threshold, or if the text recognition object is a geo-sensitive element with a volume less than or equal to a preset threshold, the first size is set to 10 pixels wide and 10 pixels high. That is, for geo-sensitive elements with a volume less than or equal to the preset threshold (such as geo-sensitive elements other than cooling towers), if the recognition box is less than or equal to 10 pixels * 10 pixels, the image recognition object is not marked as a geo-sensitive element.

[0065] Since the feature features of very small objects in the image are unclear, such objects generally cannot carry valid information. By setting this ignore box strategy, objects with unclear features can be excluded and not set as geographically sensitive features, thereby improving the accuracy of sensitive feature identification.

[0066] In some implementations, the recall strategy proposed in the embodiments of the present disclosure may also include a job reduction strategy, including at least one of the following:

[0067] (1) When the same sensitive element exists in M ​​consecutive images in chronological order, the image containing the sensitive element with the highest confidence is marked as the sensitive image; M is a positive integer.

[0068] In some embodiments, M is related to the degree of overlap between two adjacent image frames. For example, the greater the degree of overlap between two adjacent image frames, the greater the value of M.

[0069] In one example, M is greater than or equal to 3.

[0070] For example, M is 5.

[0071] For example, if five consecutive images match the same text-sensitive element / image-sensitive element, it means that there is a high possibility that sensitive information exists here. In this case, the image containing the element with the highest confidence level can be taken and marked as a sensitive image for subsequent manual work.

[0072] When the same sensitive element appears continuously in multiple images, there is a high possibility that other sensitive elements exist in the image; therefore, by adopting the above recall strategy, the image where the sensitive element appears and the confidence level of the sensitive element is the highest can be marked as a sensitive image, thereby reducing the amount of computation when identifying sensitive elements and improving recall efficiency.

[0073] (2) If the number of geographically sensitive elements in an image is greater than or equal to N, the image is marked as a sensitive image.

[0074] In one example, N is greater than or equal to 2. For example, N is 3.

[0075] For example, if an image matches multiple elements / words (such as 3 or more), it means that there is a high possibility that the image contains sensitive information. In this case, the image can be marked as a sensitive image for subsequent manual work.

[0076] When multiple sensitive elements appear in an image, it indicates that there is a greater possibility that other sensitive elements appear in the image. Therefore, by adopting the above recall strategy, the entire image can be marked as a sensitive image, rather than just a few sensitive elements in the image, thereby reducing the amount of computation required to identify sensitive elements and improving recall efficiency.

[0077] (3) When there are text-sensitive elements in an image, if the same text-sensitive elements do not exist in the X images before and after the image in the time sequence of the multiple frames, the image is considered insensitive.

[0078] In one example, X is 5.

[0079] For example, if an image matches a sensitive word in the detection vocabulary, and the sensitive word is not matched in the five images before and after the image, the image will not be marked as sensitive.

[0080] This approach can reduce the problem of mislabeling due to text recognition errors. For example, if a sensitive word appears in only one of multiple images taken consecutively, while none of the images before and after it contain the same word, it is likely that the image containing the sensitive word has been misrecognized. Therefore, the image can be marked as insensitive to correct the recognition error.

[0081] It's easy to understand that the larger the value of X, the stricter the criteria for evaluating text recognition errors. For example, a larger X means a text recognition error will be considered when a sensitive word appears only once within a longer timeframe; a smaller X means a text recognition error will be considered when a sensitive word appears only once within a shorter timeframe, thus making the evaluation criteria for text recognition errors more relaxed. The value of X can be determined based on specific circumstances. For example, if the criteria for determining sensitive information are stricter, the value of X will be larger; if the criteria for determining sensitive information are looser, the value of X will be smaller.

[0082] After the above identification, sensitive images or geographically sensitive elements in images can be identified. Then, images that are determined to contain sensitive information about units and facilities can be secondary verified to mark the image's sensitive category, sensitive area, and processing method.

[0083] In some implementations, personal sensitive information in a data packet may also be identified, for example, based on the data packet, personal sensitive information in the data packet may be identified; and the data packet may be marked according to the identified personal sensitive information.

[0084] For example, personal sensitive information includes faces, license plates, etc. The disclosed embodiment can perform personal sensitive information recognition on each frame in the data packet, thereby achieving full data packet recognition of face and license plate image features.

[0085] In some implementations, image recognition methods (e.g., image recognition models) may be used to identify sensitive personal information and return information such as the identification category, identification elements, element coordinates, and confidence level. Confidence level can be understood as the reliability of the identification result.

[0086] After identifying geographic sensitive information and personal sensitive information in a data packet, manual verification can be performed on the identified sensitive information (including at least one of geographic sensitive information and personal sensitive information) or sensitive image. Based on the results of the manual verification, the image or data packet can then be processed. Specifically, different processing methods can be used for personal sensitive information and geographic sensitive information:

[0087] (1) Processing method for personal sensitive information: Perform local contour processing on the images of recognized faces and license plate elements, and the processing results meet the anonymization requirements.

[0088] (2) Delete the images or entire task packages that are manually determined to involve geographically sensitive elements, and push the results of tasks that are determined not to involve sensitive elements.

[0089] Figure 2 FIG. 1 is a functional diagram of each processing module according to an embodiment of the present disclosure. Figure 2As shown, a data parsing model 210 is used to parse data collected by intelligent connected vehicles to obtain data packets, and then frame extraction is performed on the data packets. An image recognition module 220 is used to perform image recognition on images, and a text recognition module 230 is used to perform text recognition on images, such as using an OCR recognition method.

[0090] Specifically, the image recognition module 220 can recognize the following elements and return the recognition category, recognition element, element coordinates and confidence level.

[0091] (1) Geographically sensitive information image elements: Geographically sensitive information may include units and facilities involved in images that should be technically processed as specified in relevant regulations, including at least one of the following types:

[0092] a. Sensitive unit image features.

[0093] b. Sensitive facility image features, such as public security facility image features "oil storage depot", such as Figure 3 As shown; the image feature of civil facilities class "cooling tower", such as Figure 4 shown.

[0094] (2) Personal sensitive information image elements, including face and license plate information.

[0095] When identifying personal sensitive information, the image recognition module 220 can perform full recognition on each frame of image in the data packet; when identifying geographically sensitive information, it can recognize the image obtained after frame extraction to reduce the amount of calculation.

[0096] The text recognition module 230 can recognize text in an image and return the recognized term, its image coordinates, and its confidence level. The text recognition module 220 can use OCR technology for recognition. In some embodiments, the text recognition module 230 can perform text recognition on an image obtained after frame extraction to reduce computational complexity.

[0097] like Figure 2 As shown, the processing framework proposed in the embodiment of the present disclosure also includes a sensitive word matching module 240, which is used to determine whether the entry identified by the text recognition module 230 is a sensitive word. A matching word library can be set in the sensitive word matching module 240, and the matching word library can include a detection word library and an exemption word library. Among them, the detection word library is mainly used to identify sensitive content in images or pictures. If a match is successful, the image will be marked as sensitive.

[0098] The exemption word library is mainly used to exclude those words that appear in the image but are reasonable expressions in specific scenarios, and the image should not be marked as sensitive.

[0099] Using the above detection word library and exemption word library, the sensitive word matching module 240 can use a matching strategy to match sensitive words.

[0100] The matching strategy can be set as follows: if the word hits the detection word library and does not hit the exemption word library, it is considered sensitive.

[0101] In one example, if the matching degree between the detected word and the detection word library is greater than or equal to 75% (the threshold is configurable), and the confidence of the detected word is greater than or equal to 90% (the threshold is configurable), the detected word is marked as a sensitive type.

[0102] In another example, if the matching degree between the detected word and the exemption rule vocabulary is equal to 100% (this threshold is configurable), and the confidence of the detected word is greater than or equal to 90% (this threshold is configurable), the detected word will not be marked as a sensitive type and will proceed to the next stage.

[0103] In another example, if the matching degree between the detected word and the detection word library is less than 75% (this threshold is configurable), that is, there is no match to the content in the detection word library, the detected word will not be marked as a sensitive type and will enter the next stage.

[0104] like Figure 2 As shown, the processing framework proposed in the embodiment of the present disclosure also includes a strategy fusion module 250, which is used to determine whether the image is sensitive and the sensitive type based on the sensitive words identified by the sensitive word matching module 240 and the unit and facility image sensitive elements identified by the image recognition module 220.

[0105] In one example, a recall strategy is pre-set based on the different features of the image and OCR sensitive information. The strategy fusion module 250 can comprehensively calculate whether the image is sensitive and the sensitive category based on the recognition result and the recall strategy. The recall strategy includes at least one of the following:

[0106] 1. Recognition confidence strategy: When the confidence of the identified element is greater than or equal to the set threshold, the recognition result is considered correct.

[0107] (1) The confidence threshold of text recognition is 80%. When the confidence of text recognition is greater than or equal to 80%, the text recognition result is considered to be credible. Then, it is further determined whether the text recognition result is a sensitive word.

[0108] (2) Setting the confidence threshold for image feature recognition:

[0109] In one example, for a camouflaged house, etc., the confidence threshold is set to 80%.

[0110] For example, if an image recognition method is used to identify the presence of a camouflaged house or the like in an image, and the confidence level of the recognition result is greater than or equal to 80%, then the recognition result is considered correct, ie, a sensitive element exists in the image.

[0111] 2. Minimum detected image element strategy: When the detected feature element is too small, it means that the identification feature is not clear and it can be not marked as a sensitive element;

[0112] (1) When the recognition box is smaller than 10 pixels × 10 pixels, the element is judged as insensitive;

[0113] (2) When the identification box of a non-cooling tower or other relatively large sensitive element is smaller than 20 pixels × 60 pixels, or smaller than 60 pixels × 20 pixels, the element is determined to be insensitive.

[0114] 3. Work Reduction Strategy: Since the external video data collected by intelligent connected vehicles has a high frame rate and a high degree of image overlap, the volume of suspected sensitive images recalled based on the above model identification and strategy is very large. Therefore, a work reduction strategy can be set to reduce the workload of manual judgment:

[0115] (1) If there are multiple pictures in the current task package that are sequentially matched to the same sensitive element (including image elements and / or text elements), it means that there is a high possibility that sensitive information exists here. In this case, only the picture with the element with the highest confidence level is selected for manual work;

[0116] (2) If a picture in the current task package matches multiple sensitive elements (including image elements and / or text elements), if there are more than or equal to two sensitive elements, it means that there is a high possibility of sensitive information here, and the picture is taken for manual work;

[0117] (3) If an image matches a sensitive word in the detection vocabulary, and multiple images before and after the image (e.g., 5 images) do not match the sensitive word in chronological order, the image will not be marked as sensitive.

[0118] After the policy fusion module 250 uses the aforementioned recall strategy to mark sensitive maps or sensitive elements within them, the manual work module 260 can be used to perform a secondary verification of the content marked as sensitive maps or sensitive elements. Subsequently, the results processing module 270 combines the recall strategy with the manual verification results to ultimately determine the sensitive maps or sensitive elements. Based on this determination, the data can be deleted entirely, replaced, or deleted by packet.

[0119] The present disclosure also provides a device for processing spatiotemporal data of an intelligent connected vehicle. Figure 5 FIG. 5 is a schematic structural diagram of a spatiotemporal data processing device 500 for an intelligent connected vehicle according to an embodiment of the present disclosure, comprising:

[0120] A frame extraction module 510 is used to extract frames from data packets collected by the intelligent connected vehicle to obtain multiple frames of images, wherein the overlap between two adjacent frames of images in the multiple frames is greater than or equal to 50%;

[0121] an identification module 520 for identifying geographically sensitive information in the multiple image frames based on the multiple image frames;

[0122] The marking module 530 is configured to mark the data packet according to the recognition result.

[0123] In some implementations, a result processing module may also be included, which is used to delete, replace, or delete packets of data packets that have been identified as containing sensitive information.

[0124] In some implementations, the frame extraction method is related to the driving speed of the intelligent connected vehicle and the video frame rate of the data packet.

[0125] In some embodiments, the identification module 520 is configured to:

[0126] performing at least one of image recognition and text recognition on the multiple frames of images;

[0127] Based on a result of at least one of image recognition and text recognition and a recall strategy for recalling geographically sensitive information, geographically sensitive information in multiple frames of images is identified.

[0128] In some embodiments, the recall strategy includes at least one of the following:

[0129] If the matching degree between the result of the text recognition and the detection vocabulary is greater than or equal to the first threshold, and the matching degree between the result of the text recognition and the exemption vocabulary is less than the second threshold, the object of the text recognition is determined to be a text sensitive element;

[0130] If the result of image recognition corresponds to geographically sensitive information, the object of image recognition is determined as an image sensitive element;

[0131] When the recognition box of the image recognition is smaller than or equal to the first size, or the recognition box of the text recognition is smaller than or equal to the first size, the object of the image recognition or text recognition is determined as a non-sensitive element.

[0132] In some embodiments, the first threshold is greater than or equal to 60%.

[0133] In some embodiments, when the object of image recognition or text recognition is a geographically sensitive element with a volume greater than a preset threshold, the first size is 20 pixels in width and 60 pixels in height; or 60 pixels in width and 20 pixels in height.

[0134] In some embodiments, when the object of image recognition is a geographically sensitive element with a volume less than or equal to a preset threshold, or the object of text recognition is a geographically sensitive element with a volume less than or equal to a preset threshold, the first size is 10 pixels in width and 10 pixels in height.

[0135] In some embodiments, the recall strategy further comprises:

[0136] When the same sensitive element exists in M ​​images consecutively in chronological order, the image containing the geographically sensitive element with the highest confidence is marked as a sensitive image; M is a positive integer.

[0137] In some embodiments, M is related to the overlap between two adjacent image frames.

[0138] In some embodiments, the recall strategy further comprises:

[0139] If the number of geographically sensitive elements in an image is greater than or equal to N, the image is marked as a sensitive image, where N is a positive integer.

[0140] In some embodiments, the recall strategy further comprises:

[0141] When there are text-sensitive elements in an image, according to the time sequence of multiple frames, if the same text-sensitive elements do not exist in the X images before and after the image, the image is considered insensitive, where X is a positive integer.

[0142] In some embodiments, the identification module 520 is further used to identify personal sensitive information in the data packet based on the data packet; the marking module 530 is further used to mark the data packet according to the identified personal sensitive information.

[0143] For the description of specific functions and examples of each module and submodule of the device in the embodiment of the present disclosure, please refer to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.

[0144] In the technical solution disclosed herein, the acquisition, storage and application of personal information of users involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0145] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0146] Figure 6A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0147] like Figure 6 As shown, the device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the device 600 can also be stored in the RAM 603. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0148] Multiple components in device 600 are connected to I / O interface 605, including: input unit 606, such as a keyboard, mouse, etc.; output unit 607, such as various types of displays, speakers, etc.; storage unit 608, such as a magnetic disk, optical disk, etc.; and communication unit 609, such as a network card, modem, wireless communication transceiver, etc. Communication unit 609 allows device 600 to exchange data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0149] The computing unit 601 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 501 performs the various methods and processes described above, such as the detection method. For example, in some embodiments, the detection method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the detection method described above can be performed. Alternatively, in other embodiments, the computing unit 601 can be configured to perform the detection method in any other appropriate manner (e.g., by means of firmware).

[0150] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0151] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0152] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0153] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0154] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0155] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0156] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0157] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A method for processing spatiotemporal data of an intelligent connected vehicle, comprising: Extracting frames from data packets collected by the intelligent connected vehicle to obtain multiple frames of images, wherein the overlap between two adjacent frames of the multiple frames is greater than or equal to 50%; Based on the multiple frames of images, identifying geographically sensitive information in the multiple frames of images; The data packet is marked according to the recognition result.

2. The method according to claim 1, wherein The frame extraction method is related to the driving speed of the intelligent connected vehicle and the video frame rate of the data packet.

3. The method according to claim 2, wherein: The identifying, based on the multiple frames of images, geographically sensitive information in the multiple frames of images includes: performing at least one of image recognition and text recognition on the multiple frames of images; Based on a result of at least one of the image recognition and the text recognition and a recall strategy for recalling geographically sensitive information, geographically sensitive information in the multiple frames of image is identified.

4. The method according to claim 3, wherein: The recall strategy includes at least one of the following: If the matching degree between the result of the text recognition and the detection word library is greater than or equal to a first threshold, and the matching degree between the result of the text recognition and the exemption word library is less than a second threshold, determining the object of the text recognition as a text sensitive element; If the result of the image recognition corresponds to the geographically sensitive information, determining the object of the image recognition as an image sensitive element; When the recognition box of the image recognition is smaller than or equal to the first size, or the recognition box of the text recognition is smaller than or equal to the first size, the object of the image recognition or text recognition is determined as a non-sensitive element.

5. The method according to claim 4, wherein The first threshold is greater than or equal to 60%.

6. The method according to claim 4, wherein: When the object of the image recognition or text recognition is a geographically sensitive element whose volume is larger than a preset threshold, the first size is 20 pixels in width and 60 pixels in height; or 60 pixels in width and 20 pixels in height.

7. The method according to claim 4, wherein: When the object of image recognition is a geographically sensitive element whose volume is less than or equal to a preset threshold, or the object of text recognition is a geographically sensitive element whose volume is less than or equal to a preset threshold, the first size is 10 pixels in width and 10 pixels in height.

8. The method according to any one of claims 4 to 7, wherein: The recall strategy also includes: When the same geographically sensitive element exists in M ​​consecutive images in chronological order, the image containing the geographically sensitive element with the highest confidence is marked as a sensitive image; M is a positive integer.

9. The method according to claim 8, wherein The M is related to the overlap degree between the two adjacent frames of images.

10. The method according to any one of claims 4 to 7, wherein: The recall strategy also includes: When the number of geographically sensitive elements in the image is greater than or equal to N, the image is marked as a sensitive image, where N is a positive integer.

11. The method according to any one of claims 4 to 7, wherein: The recall strategy also includes: In the case where text-sensitive elements exist in the image, according to the time sequence of the multiple frames of images, if the same text-sensitive elements do not exist in the X images before and after the image, the image is marked as insensitive, where X is a positive integer.

12. The method according to any one of claims 1 to 7, further comprising: Based on the data packet, identifying personal sensitive information in the data packet; The data packet is marked according to the identified personal sensitive information.

13. A spatiotemporal data processing device for an intelligent connected vehicle, comprising: A frame extraction module is used to extract frames from data packets collected by the intelligent connected vehicle to obtain multiple frames of images, wherein the overlap between two adjacent frames of images in the multiple frames is greater than or equal to 50%; an identification module, configured to identify geographically sensitive information in the multiple image frames based on the multiple image frames; The marking module is used to mark the data packet according to the recognition result.

14. The device according to claim 13, wherein The frame extraction method is related to the driving speed of the intelligent connected vehicle and the video frame rate of the data packet.

15. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 12.

16. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-12.

17. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Semanteme based geographical label content safe checking method and device

    CN104008169A

  • Video frame extraction method and system based on deep learning

    CN113792600A

  • Method for automatically detecting and desensitizing sensitive information in picture acquired based on high-precision map

    CN114463755A

  • Automatic desensitization processing method and system for sensitive information in road traffic video

    CN115810160A

  • Data processing method and device and storage medium

    CN116861198A