A method for inducing the release of content two-factor detection authentication

By performing consistency verification and anomaly detection on the pre-release processing data and original data of the inducement release device, and combining it with post-release image comparison, the problem of insufficient content correction of the inducement release device was solved, realizing full-process control and improving the accuracy and security of information release.

CN121545142BActive Publication Date: 2026-04-28CHENGDU JIAOTOU INTELLIGENT TRANSPORTATION TECHNOLOGY SERVICE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHENGDU JIAOTOU INTELLIGENT TRANSPORTATION TECHNOLOGY SERVICE CO LTD
Filing Date
2026-01-20
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In existing technologies, the content published by misleading publishing devices cannot be effectively corrected before publication, which increases the risk of misleading information. The existing review method is post-event verification, which cannot solve the problem at the source.

Method used

By employing consistency verification and anomaly detection between pre-release processed data and original data, combined with comparison of the displayed image after release with the pre-release processed data, a full-process control system is formed to ensure content consistency and implement emergency strategies in case of anomalies.

Benefits of technology

It eliminates data processing deviations at the source, intercepts abnormal content, and improves the accuracy and security of misleading content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121545142B_ABST
    Figure CN121545142B_ABST
Patent Text Reader

Abstract

The application provides a kind of induced release content two-factor detection authentication method, comprising: obtaining pre-release processing data and pre-release original data, pre-release processing data represents the data stored in the database after pre-release original data is processed according to format requirements;Determine content consistency based on pre-release processing data and pre-release original data;When the content is consistent, perform anomaly detection on pre-release processing data;When there is no anomaly, send pre-release processing data to the induced release device for display;Collect induced release device display image;Determine emergency strategy based on induced release device display image and corresponding pre-release processing data. By implementing the application, the accuracy and security of induced release content are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of information release management technology, specifically relating to a two-factor detection and authentication method for inducing content to be released. Background Technology

[0002] In fields such as intelligent transportation, commercial advertising, and public information dissemination, information guidance devices, such as outdoor LED screens and traffic sign screens, are key carriers of information transmission. The accuracy and security of their displayed content directly affect public decision-making and public order. Regarding the inspection of the content displayed by these devices, current technologies typically use cameras to review the displayed content. However, this is a post-event review and does not address the problem at its source. Errors in the content displayed by these devices cannot be corrected before publication, increasing the risk of misleading information being disseminated. Summary of the Invention

[0003] In view of this, the purpose of this invention is to provide a two-factor authentication method for inducing the publication of content, so as to meet the needs of reducing the risk of publishing misleading information and improving the accuracy of information publication.

[0004] To achieve the above objectives, the present invention provides the following technical solution:

[0005] According to a first aspect, the present invention provides a two-factor authentication method for inducing content publication, comprising: acquiring pre-publication processing data and pre-publication original data, wherein the pre-publication processing data represents the data stored in a database after the pre-publication original data has been processed according to format requirements; determining content consistency based on the pre-publication processing data and the pre-publication original data; if the content is consistent, performing anomaly detection on the pre-publication processing data; if there is no anomaly, sending the pre-publication processing data to an inducing publication device for display by the inducing publication device; acquiring the image displayed by the inducing publication device; and determining an emergency strategy based on the image displayed by the inducing publication device and the corresponding pre-publication processing data.

[0006] According to the second aspect, a two-factor authentication device for inducing content publication includes: a data acquisition module for acquiring pre-publication processed data and pre-publication original data, wherein the pre-publication processed data represents the data stored in a database after the pre-publication original data has been processed according to format requirements; a consistency determination module for determining content consistency based on the pre-publication processed data and the pre-publication original data; an anomaly detection module for performing anomaly detection on the pre-publication processed data when the content is consistent; a sending module for sending the pre-publication processed data to an inducing publication device when there is no anomaly, so that the inducing publication device can display it; a display module for acquiring the image displayed by the inducing publication device; and an emergency strategy determination module for determining an emergency strategy based on the image displayed by the inducing publication device and the corresponding pre-publication processed data.

[0007] According to a third aspect, an embodiment of the present invention provides an electronic device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor performs the steps of the two-factor authentication method for induced content publishing described in the first aspect or any embodiment of the first aspect.

[0008] According to a fourth aspect, embodiments of the present invention provide a computer storage medium storing computer instructions thereon, which, when executed by a processor, implement the steps of the two-factor detection and authentication method for induced content publication as described in the first aspect or any embodiment of the first aspect.

[0009] This embodiment provides a two-factor authentication method for inducing content to be published. It forms a two-factor detection before publication by verifying the consistency between pre-published processed data and the original data and detecting anomalies in the pre-published processed data. Combined with the comparison between the displayed image after publication and the pre-published processed data, it achieves full-process control, eliminates data processing deviations from the source, intercepts abnormal content during the process, and monitors display risks from the terminal, thereby improving the accuracy and security of inducing content to be published.

[0010] Other advantages, objectives, and features of the invention will be set forth in the following description and will be apparent to those skilled in the art in some respects, or may be learned by practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description

[0011] To make the objectives, technical solutions, and beneficial effects of this invention clearer, the following figures are provided for illustration:

[0012] Figure 1 This is a flowchart illustrating a specific example of a two-factor authentication method for inducing content publication according to the present invention.

[0013] Figure 2 This is a specific modular example diagram of a two-factor authentication device for inducing content publication in this invention;

[0014] Figure 3 This is a schematic block diagram of a specific example of an electronic device in an embodiment of the present invention. Detailed Implementation

[0015] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0016] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can also refer to the internal connection of two components; and they can refer to a wireless connection or a wired connection. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0017] Furthermore, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0018] Currently, various organizations have strict requirements for information security management of LED display equipment. For outdoor induction display equipment (LED screens), a dual mechanism of "pre-display verification + post-display monitoring" is used to ensure information dissemination security. Therefore, this invention provides a two-factor detection and authentication method for induction display content, such as... Figure 1 As shown, it includes:

[0019] S101, Obtain pre-release processing data and pre-release raw data. The pre-release processing data represents the data stored in the database after the pre-release raw data has been processed according to the format requirements.

[0020] S102, Based on the pre-release processed data and the pre-release original data, determine the consistency of content;

[0021] S103, If the content is consistent, perform anomaly detection on the pre-release processing data;

[0022] S104, if there are no abnormalities, the pre-release processing data is sent to the inducing release device so that the inducing release device can display it;

[0023] S105, Collect images displayed by the guidance and release device;

[0024] S106, Based on the image displayed by the induced release device and the corresponding pre-release processing data, determine the emergency strategy.

[0025] For example, this embodiment uses the field of intelligent transportation as an example for illustration. The guidance and release device can be an electronic screen installed outdoors. The pre-release processing data can be JPG format data generated based on the pre-release original data. The pre-release original data can be initial data that has not been formatted and conforms to the source attributes of the business scenario, such as vector graphics original data stored in SVG format or design source files in PSD format. The pre-release processing data can be obtained from a Redis database. At the same time, in order to ensure the security of data storage, the pre-release processing data can be encrypted data. After the encrypted data is extracted from Redis, it is restored to an image. Then, based on the pre-release processing data and the pre-release original data, content consistency detection is implemented. The detection process is as follows:

[0026] As an optional implementation, when the pre-released original data contains graphic and textual information, the content consistency detection process includes: parsing the pre-released original data structure and extracting the first visual element, which includes the position and size of the shape, the content and position of the text, the RGB values ​​of the color, and the element level; identifying the pre-released processed data based on an image recognition algorithm and extracting the second visual element, which includes the position and size of the shape, the content and position of the text, the RGB values ​​of the color, and the element level; and performing a quantitative comparison between the first visual element and the second visual element to obtain the consistency verification result, which includes the presence of the element, positional deviation, size deviation, text matching degree, RGB, and element level matching degree.

[0027] Specifically, if the pre-released raw data is a design source file, such as PSD format, the development interface of the corresponding design software is called to parse the file's layer structure, object attributes, and metadata. The hierarchical relationship of each visual element is obtained by traversing the layers, with the top layer denoted as "Level 1," and the layers numbered sequentially downwards. For shape elements, a pixel coordinate system is established with the top-left corner of the data as the origin, and the anchor point coordinates and the width and height dimensions of the bounding rectangle are extracted. For text information, the text content is extracted. For color elements, the RGB values ​​of the primary color of each visual element are collected. If the pre-released raw data is a structured image, such as SVG format, the vector path information of the file is parsed using an XML parser to identify the path coordinates and fill color RGB values ​​of the shape elements, and the size and position of the bounding rectangle of the shape are calculated. The text content is parsed through text tags, and the hierarchy of each visual element is determined based on the path hierarchy. After extraction, the first visual element is stored in the form of a structured dataset, with each data entry containing fields such as element type, position coordinates, size, RGB value, and hierarchy number.

[0028] If the pre-release processing data is in image format, the extraction of second visual elements is achieved based on image recognition algorithms. Specifically, object detection, OCR, and color analysis are used to extract visual elements. The YOLOv5 object detection model detects the position of each visual element, outputting the coordinates and dimensions of the element's bounding rectangle. For detected text regions, an OCR engine combined with an LSTM recognition model is used to extract the text content. For each detected visual element, the average RGB value within its bounding rectangle is calculated as the element's primary color RGB value. Layering is determined by element occlusion relationships. Specifically, for any two elements A and B, if the bounding rectangles of A and B overlap, and the pixel percentage of A within the overlapping area exceeds 80%, then A is considered to be at a higher layer than B. The extracted second visual elements are organized according to the structured dataset format of the first visual elements to ensure one-to-one field correspondence for subsequent quantitative comparison. The OCR, YOLOv5, and LSTM recognition models used above for text extraction are existing technologies and will not be elaborated upon here.

[0029] The specific process for evaluating the consistency dimension of element existence is as follows: Count the number of missing elements in the first visual element that do not match the second visual element, and the number of newly added elements in the second visual element that do not match the first visual element. If both the number of missing elements and the number of newly added elements are 0, then the consistency dimension evaluation passes.

[0030] The specific process for evaluating the consistency dimension of positional deviation is as follows: Calculate the deviation between the positional coordinates of the second visual element and the corresponding coordinates in the first visual element. The deviation value is calculated using the Euclidean distance formula. If the deviation value... If the deviation is less than 10 pixels, the consistency dimension is evaluated as passing; if it exceeds 10 pixels, it is marked as an abnormal positional deviation and the specific deviation value is recorded.

[0031] The specific process for evaluating the consistency of size deviation is as follows: calculate the percentage deviation of the width and height dimensions of the elements in the second visual element from the corresponding dimensions in the first visual element. If the percentage deviations of the width and height dimensions are both... If the deviation is 3%, the consistency dimension is considered to have passed the evaluation; if any dimension exceeds 3%, it is marked as an abnormal size deviation.

[0032] The specific process for evaluating the consistency dimension of RGB matching is as follows: For text information, it is determined whether the text in the second visual element is completely consistent with the text in the first visual element. If they are completely consistent, the consistency dimension evaluation is passed.

[0033] The specific process for evaluating the consistency of color and hierarchy is as follows: If the deviations of the R, G, and B channels of the RGB values ​​are all... If the 10th level and the level number are consistent, then the consistency dimension assessment is passed; otherwise, it is marked as a color anomaly or a level anomaly.

[0034] If any one of the above five consistency verification aspects fails, the consistency verification result is considered unsuccessful. If all five aspects pass, the content is considered consistent, and the consistency verification result is considered successful.

[0035] As another optional implementation, when the pre-released original data itself only contains text information, the content consistency detection process includes: using an image recognition algorithm to recognize the image and obtain the text information in the image; inputting the pre-released original data into a pre-trained first intent recognition model to obtain the intent of the pre-released original data; inputting the text information into a pre-trained second intent recognition model to obtain the intent of the pre-released processed data; and matching the intent of the pre-released original data with the intent of the pre-released processed data to determine content consistency.

[0036] For example, a combination of object detection and OCR recognition is adopted. First, the image is scanned by the YOLOv5 object detection model to locate the bounding rectangle coordinates of all text regions and exclude non-text interference areas. Then, the OCR engine equipped with LSTM is called to perform character recognition on each text region and output complete text information.

[0037] The structured information in the pre-released raw data is converted into a text sequence and input into a pre-trained first intent recognition model. The first intent recognition model can be a BERT-based intent recognition model. The model outputs the core intent of the pre-released raw data by semantically encoding the text sequence and generates an intent vector.

[0038] The complete text information obtained through image recognition algorithm is input into the pre-trained second intent recognition model. The second intent recognition model can have the same structure and parameters as the first intent recognition model. That is, the second intent recognition model can also be a BERT-based intent recognition model.

[0039] The BERT-based intent recognition model is an existing model with the BERT-based pre-trained model as its core architecture. Its foundation is a 12-layer stacked Transformer encoder, containing 12 attention heads and 768-dimensional hidden layers. Each Transformer encoder layer consists of a multi-head self-attention mechanism and a feedforward neural network. The former calculates attention scores and fuses multi-view semantics by projecting the input vector into query, key, and value vectors. The latter enhances features through a linear transformation of "768-dimensional → 3072-dimensional → 768-dimensional" and ReLU activation. Both layers use residual connections and layer normalization for stable training. The input is processed by WordPiece word segmentation, adding [CLS] and [SEP] tags, and combining word embeddings, position embeddings, and sentence embeddings to form the input vector. At output, the 768-dimensional context representation of the [CLS] tag is extracted and mapped to intent category probabilities through a classification head. The entire process is completed through pre-training and fine-tuning to achieve accurate recognition of text intent.

[0040] Finally, calculate the cosine similarity between the two intent vectors, set a similarity threshold, and if the similarity is... If the threshold is met and the textual descriptions of the two core intents do not conflict, then the content is considered to be consistent. For example, if the original intent is "speed limit 60", and the processed intent is "speed limit 60km / h", then it is considered to be without conflict.

[0041] If the consistency verification passes, anomaly detection is performed. This anomaly detection may involve sensitive word screening, where sensitive words can be predefined and stored in a database. If no sensitive words are found, the pre-released data is sent to the inducing release device for display. Surveillance cameras are deployed on the inducing release device, and the video signal is connected to the backend system. Using video analysis algorithms, the actual content released by the device is compared in real time with the program content sent from the backend to ensure consistency. If inconsistency is found, an emergency strategy is activated. This strategy may include prioritizing the "revoke release" or "overwrite release" operation on the inducing device, or remotely powering off the device via its built-in remote power-off switch to completely stop the spread of illegal content.

[0042] This invention provides a two-factor authentication method for inducing content to be published. It forms a two-factor detection before publication by verifying the consistency between pre-published processed data and the original data and detecting anomalies in the pre-published processed data. Combined with the comparison between the displayed image after publication and the pre-published processed data, it achieves full-process control, eliminates data processing deviations from the source, intercepts abnormal content during the process, and monitors display risks from the terminal, thereby improving the accuracy and security of inducing content to be published.

[0043] As an optional implementation, the pre-release processing data is a combination of text and image data. Before determining content consistency based on the pre-release processing data and the original pre-release data, the process includes: performing image-to-text segmentation on the pre-release processing data to obtain target text and target image; extracting text focus from the target text based on a pre-trained target language model, where text focus includes the text subject, action, and condition; determining visual focus anchors in the target image based on a pre-trained target detection model, where visual focus anchors include target anchors, action anchors, and numerical anchors; inputting the subject in the target text and the corresponding target anchors in the target image into a pre-trained target machine learning model to determine the vector similarity between the subject and the target anchors in the semantic space; comparing the action in the target text with the action represented by the corresponding action anchor in the target image to calculate the action matching degree; parsing the condition in the target text into a logical expression and substituting the numerical value represented by the corresponding numerical anchor in the target image into the corresponding logical expression to determine the numerical matching degree; and determining whether to execute the step of determining content consistency based on the pre-release processing data and the original pre-release data based on the vector similarity, action matching degree, and numerical matching degree.

[0044] For example, this embodiment is applicable to scenarios where the pre-released processed data is a combination of text and images. By verifying the semantic association between text and images, the pre-released processed data is checked before the data is displayed on the inducing publishing device, and data with obvious semantic mismatches is discovered in advance to avoid ambiguity to the public caused by the displayed data.

[0045] In this embodiment, a semantic segmentation model based on an improved U-Net is selected. This model enhances the segmentation capability of small text regions and text regions in complex backgrounds by introducing nested densely connected blocks. The model input is the original format of combined image and text data, such as JPG format. Before input, the image size needs to be unified and the pixel values ​​need to be normalized to the [0, 1] interval to reduce computational complexity.

[0046] The preprocessed image and text data is input into the trained U-Net model. The model outputs a binary segmentation mask with the same size as the input, where the pixel value for text regions is 1 and the pixel value for image regions is 0. The model training phase can cover multiple typical image and text scenarios, such as traffic, parking, commercial advertisements, and equipment instructions, to ensure segmentation generalization.

[0047] Based on segmentation masks, ROI cropping technology is used to separate text and images. For text regions, connected regions with a pixel value of 1 are extracted from the mask, located by an outer rectangle, and cropped to obtain a grayscale image containing only text. At the same time, the coordinate information of this region in the original text and image data is recorded. For image regions, regions with a pixel value of 0 are extracted from the mask, and cropped to obtain pure image data without text interference, while retaining the original RGB color information.

[0048] Next, the pre-trained RoBERTa-wwm-ext-large Chinese model was selected as the target language model. This model, through full-word masking and expanded training corpus, performs better in Chinese semantic understanding and entity recognition tasks, accurately extracting subject, action, and conditional information from the text. The specific process is as follows:

[0049] First, the segmented text regions are converted into editable text using OCR technology, followed by text cleaning to form a standardized text sequence. This standardized text sequence is then input into the RoBERTa model, where the model's NER layer identifies and labels core entities within the text, initially filtering out potential text subjects and numerical conditions. Examples of text subjects include: "Left lane parking lot area A," and numerical conditions include: 7:00-9:00, speed limit 50km / h, 3 empty parking spaces.

[0050] For the action extraction process, this embodiment analyzes the semantic relationships between entities based on the syntactic dependency tree output by the model, determines the text actions and action directions, and filters out verbs / verb phrases that represent states or behaviors as the actions to be extracted. For the condition extraction, this embodiment determines the correspondence between the action and the subject through subject-predicate and verb-object relationships, and determines the correspondence between the action and the condition through adverbial-head relationships.

[0051] Finally, the extracted results are organized to form a structured text focus list, clarifying the correspondence between the subject, action, and condition. For example, if the text content is: "During the morning rush hour from 7:00 to 9:00, the left lane of XX Road is closed, and passing vehicles are requested to detour via the right lane," then the subjects in the text focus are: the left lane of XX Road, passing vehicles, and the right lane; the actions are: closure and detour; and the condition is: morning rush hour from 7:00 to 9:00.

[0052] Next, the Faster R-CNN model was used as the object detection model, which utilizes RPN. FastR-CNN's two-stage structure can accurately locate multiple types of targets in images and output the target's category and location information, making it suitable for extracting visual focus anchors. The model has been fine-tuned for common scenarios involving combined image and text data, supporting the recognition of three types of anchors: target anchors, action anchors, and numerical anchors. The specific process is as follows:

[0053] First, the segmented image data is adjusted to the model input format, preserving key visual features such as color and shape. Potential target regions are generated using RPN, and then FastR-CNN is used to classify and regress bounding boxes in these regions, outputting detailed information for three types of anchors: category, confidence level, and bounding rectangle coordinates. Target anchors are defined as visual entities corresponding to the main text, such as lane areas, vehicle icons, parking lot partitions, and equipment components, with category labels such as lane, vehicle, and partition. Action anchors are defined as visual symbols representing action states, such as a red cross for closure and a green arrow for detour, with category labels such as closure symbol and detour arrow. Numerical anchors are defined as numbers and unit labels in the image, such as 7:00, 50, 3, and km / h, with category labels such as time, speed, quantity, and unit.

[0054] Finally, based on all the output visual focus anchors, a list of visual focus anchors is generated, specifying the anchor type, content, and location coordinates. For example, an example of an image anchor list is as follows:

[0055] Target anchor points include:

[0056] The coordinates of the left lane area of ​​XX Road are: (100, 200) - (300, 400);

[0057] The coordinates of the right lane area are: (400, 200) - (600, 400);

[0058] The vehicle icon has coordinates (250, 350) - (280, 380).

[0059] Action anchor points include:

[0060] The red cross symbol on the left lane has coordinates of (180, 280) - (220, 320) and is classified as a closed symbol.

[0061] The green right-pointing arrow in the right lane has coordinates of (480, 280) - (520, 320) and is categorized as a detour arrow.

[0062] Numerical anchor point:

[0063] The time period is 7:00-9:00, and its coordinates are: (50, 50)-(150, 80);

[0064] Finally, based on the extracted text focus and visual focus anchor points, the semantic relationship between the two is comprehensively verified through vector similarity, action matching degree, and numerical conformity. Thresholds are set for each type of verification to ensure the objectivity and reliability of the verification results.

[0065] First, the vector similarity between the subject and the target anchor point is calculated as follows: the semantic information of the text subject and the visual information of the target anchor point are mapped to the same semantic space, and the degree of matching is quantified by vector similarity. The Siamese network (two-branch structure) is selected as the target machine learning model, which is good at measuring the similarity between two samples. The specific operation is as follows:

[0066] Each subject in the text focus (e.g., the left lane of XX Road) is input into a pre-trained RoBERTa model, and the output of the last layer is extracted as a semantic vector. The corresponding target anchor point region in the visual focus anchor point is input into a pre-trained ResNet-50 model, and the output of the global average pooling layer is extracted as a visual vector. Then, a fully connected layer maps it to the text vector dimension to unify it. The text subject vector and the target anchor point vector are input into a Siamese network, and the similarity value between the two is calculated through the network's cosine similarity layer.

[0067] A similarity threshold of 0.7 is set. If the similarity is ≥0.7, the subject and the target anchor point are considered semantically matched; if it is <0.7, they are considered mismatched. For example, if the text vector of the left lane of XX Road has a similarity of 0.82 with the visual vector of the left lane area, then it matches; if it has a similarity of 0.21 with the visual vector of the right lane area, then it does not match.

[0068] For text actions, the first step is to identify scenarios where the text action contains directional information. When directional information is included, it is encoded into a first direction vector based on a predefined encoding rule. The direction of the corresponding action anchor point in the target image is determined based on an image gradient algorithm. The direction of the action anchor point is encoded based on the predefined encoding rule to obtain a second direction vector. The dot product result is determined based on the first and second direction vectors. The action matching degree is determined based on the dot product result.

[0069] For example, the direction is determined by combining natural language understanding and encoded as a unit vector as the first direction vector. The predefined encoding rule can be that the origin is the top left corner of the image, the positive x-axis is horizontal to the right, and the positive y-axis is vertical downward. That is, when driving around the right side, the direction vector is (1, 0); when driving around the left side, the direction vector is (-1, 0); when driving upward, the direction vector is (0, -1); and when driving downward, the direction vector is (0, 1).

[0070] For action anchors with directional information, such as arrows, the arrow direction is determined through image gradient analysis. Following the same encoding rules as the first direction vector (i.e., with the top-left corner of the image as the origin, horizontal to the right as the positive x-axis, and vertical downwards as the positive y-axis), the arrow direction is converted into a corresponding unit vector. The dot product of the text action direction vector and the action anchor direction vector is calculated, with a dot product threshold of 0.8. A score of 0.8 indicates a satisfactory motion matching score.

[0071] Determining the arrow direction through image gradient analysis includes: cropping a rectangular Region of Interest (ROI) containing only the arrow from the target image based on the detected coordinates of the arrow-type action anchor point; converting the ROI region into a single-channel grayscale image to obtain a first image; enhancing the contrast of the first image through adaptive histogram equalization to obtain a second image; processing the second image using a thresholding algorithm to obtain a binary image; convolving the binary image using the Sobel operator to obtain the directional gradient matrix; determining the gradient magnitude and gradient direction angle of each pixel based on the pixel magnitude in the first image; determining the effective pixels based on a pre-set gradient magnitude threshold and the pixel gradient magnitude; dividing the gradient direction angle of the effective pixels into multiple directional partitions; calculating the effective pixel gradient magnitude and the weighted sum of pixel values ​​within each directional interval; and using the directional interval containing the maximum weighted sum as the dominant directional interval to obtain the direction of the action anchor point.

[0072] Specifically, a rectangular Region of Interest (ROI) containing only arrows is cropped from the target image. The ROI region is then converted into a single-channel grayscale image and denoised using Gaussian blur. Next, adaptive histogram equalization is applied to enhance the contrast of the ROI region. Then, Otsu's automatic thresholding method is used to convert it into a binary image, making the arrows the foreground and the background black. Finally, Sobel operators in the x and y directions are used to convolve the preprocessed binary ROI image to obtain the gradient matrix in the x direction. and the gradient matrix in the y-direction For each pixel in the ROI region ( i,j That is, each pixel in the first image (single-channel grayscale image) i,j ), calculate gradient magnitude and gradient direction angle .

[0073] Next, set an amplitude threshold T and retain valid pixels with gradient amplitudes greater than T; then set the gradient direction angle of the valid pixels. Mapped to [0°, 360°], the area is divided into 8 directional intervals at 45° intervals. The weighted sum of the magnitudes of all valid pixels within each directional interval containing valid pixels is calculated, with the weight being the gradient magnitude of the corresponding pixel. The interval with the largest weighted sum of magnitudes is selected as the dominant directional interval. If the dominant directional interval is 0°-45° or 315°-360°, the arrow points to the "right," corresponding to the unit vector (1,0); if the dominant directional interval is 135°-225°, the arrow points to the "left," corresponding to the unit vector (-1,0); if the dominant directional interval is 45°-135°, the arrow points to the "up," corresponding to the unit vector (0,-1); if the dominant directional interval is 225°-315°, the arrow points to the "down," corresponding to the unit vector (0,1).

[0074] For state-type action anchors (such as red crosses and green checkmarks), that is, scenarios that do not contain directional information, their category labels are directly obtained. The state-type action is directly matched with the action anchor category. For example, if closed corresponds to the closed symbol, it passes; if open corresponds to the prohibited symbol, it fails.

[0075] For conditional information in the text focus, semantic parsing is used to convert it into standardized logical expressions, supporting logical relationships such as equal to, greater than, and less than: For example, the logical expression for morning rush hour (7:00-9:00) is: time ∈ [7:00, 9:00]; the logical expression for a speed limit of 50 km / h is: speed 50km / h; Parking lot A has 3 empty spaces, the logical expression is: number of empty spaces = 3. Extract the numerical values ​​and unit information from the visual focus anchor points, and standardize them according to the format of the text conditions to ensure that the numerical format is consistent with the text conditions. Substitute the standardized visual values ​​into the logical expression. If the expression is satisfied, the numerical compliance is passed; otherwise, it is not passed. For example, if the text condition expression is time ∈ [7:00, 9:00], and the visual value is 7:00-9:00, it satisfies the expression and passes; if the visual value is 10:00-12:00, it does not satisfy the expression and fails.

[0076] The verification results are based on three dimensions: vector similarity, action matching degree, and numerical conformity. A judgment logic of allowing all to pass and blocking any one to fail is adopted to ensure that the combined image and text data entering the subsequent detection has basic semantic consistency.

[0077] This invention provides a two-factor authentication method for inducing content publication. For combined text and image data, it achieves deep correlation verification between text and images through text focus extraction, visual focus anchor point recognition, and multi-dimensional semantic comparison, avoiding information ambiguity caused by the disconnect between text and images and improving the security of publishing text and image data.

[0078] Based on the above process, a fixed gradient direction angle division rule is adopted, uniformly dividing the area into 8 intervals, each interval being 45°. For slender arrows, the gradient directions they point to have small differences and are highly concentrated. A fixed coarse-grained division will cause gradients with different fine directions to be grouped into the same interval, making it impossible to distinguish between neighboring directions such as 30° and 45°, resulting in direction detection bias. For thick short arrows, the gradient directions are distributed over a wide range. A fixed fine-grained division will cause the effective gradients of the same arrow to be dispersed into multiple intervals, making it impossible to form a clear weighted and maximum value interval, which is prone to misjudging the dominant direction.

[0079] Therefore, as an optional implementation, the gradient direction angles of the effective pixels are divided to obtain multiple directional partitions, including: segmenting the arrow region from the binary image based on an image segmentation algorithm; calculating the width and length values ​​of the arrow region; selecting the maximum width value as the target width value and the maximum length value as the target length value; determining the number of directional partitions based on the target length value and the target width value; and dividing the gradient direction angles of the effective pixels based on the number of directional partitions to obtain multiple directional partitions.

[0080] For example, firstly, the binary image undergoes morphological preprocessing. First, erosion is used to eliminate isolated noise points smaller than a preset size structural element. Then, dilation is used to restore the pixel connectivity of the arrow's main outline. Next, connected regions whose area meets a preset minimum arrow area threshold and which have a small percentage of background residue are selected. If adjacent connected regions exist and are close together, they are merged to obtain a complete arrow region. Then, the outer contour of the arrow region is extracted and fitted with a minimum bounding rectangle. Simultaneously, the single-pixel skeleton of the arrow region is obtained. The target length value is determined by combining the length of the long side of the bounding rectangle with the arc length of the skeleton. The distance between the intersection points of the arrow contours along the vertical direction of the skeleton is calculated, and the maximum value is taken as the target width value. The obtained target length and width values ​​are then compared... The process involves standardizing pixel units and filtering outliers. Then, the length-to-width ratio is calculated. Based on the specific length and width values ​​of the target, the number of directional partitions is determined according to preset rules. For example, a length-to-width ratio ≥ 5 and a length ≥ 100px constitutes an ultra-thin, long arrow, corresponding to 12 partitions; a length-to-width ratio ≤ 0.8 and < 1.5 and a width ≥ 20px constitutes a thick, short arrow, corresponding to 6 partitions, etc. The previously obtained effective pixel gradient direction angles are then mapped from -180° to 180° to 0° to 360°. Starting from 0°, continuous, non-overlapping direction intervals are sequentially divided according to the angle span of 360° divided by the number of partitions. Finally, all effective pixels are traversed, and their gradient direction angles are assigned to the corresponding intervals, completing the partitioning and classification of effective pixel gradient directions. Ultimately, this ensures that arrows of different scales can match the optimal direction division rules, significantly improving the accuracy of direction detection.

[0081] This embodiment provides a two-factor authentication device for inducing content posting, such as... Figure 2 As shown, it includes:

[0082] Data acquisition module 201 is used to acquire pre-release processed data and pre-release raw data. The pre-release processed data represents the data stored in the database after the pre-release raw data has been processed according to the format requirements.

[0083] The consistency determination module 202 is used to determine content consistency based on the pre-release processed data and the pre-release original data;

[0084] The anomaly detection module 203 is used to perform anomaly detection on the pre-release processing data when the content is consistent.

[0085] The sending module 204 is used to send the pre-release processing data to the inducing release device when there is no abnormality, so that the inducing release device can display it;

[0086] Display module 205 is used to collect images displayed by the inducement release device;

[0087] The emergency strategy determination module 206 is used to determine the emergency strategy based on the image displayed by the induced release device and the corresponding pre-release processing data.

[0088] This application also provides an electronic device, such as... Figure 3 As shown, processor 501 and memory 502 are connected via a bus or other means.

[0089] Processor 501 can be a central processing unit (CPU). Processor 501 can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or combinations of the above types of chips.

[0090] The memory 502, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the two-factor authentication method for induced content publishing in this embodiment of the invention. The processor executes various functional applications and data processing by running the non-transitory software programs, instructions, and modules stored in the memory.

[0091] Memory 502 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created by the processor, etc. Furthermore, the memory may include high-speed random access memory and non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory 502 may optionally include memory remotely located relative to the processor, which can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0092] The one or more modules are stored in the memory 502, and when executed by the processor 501, they perform actions such as... Figure 1 The illustrated embodiment presents a two-factor authentication method for inducing content publication.

[0093] For specific details regarding the aforementioned electronic devices, please refer to the relevant documentation. Figure 1 The relevant descriptions and effects in the illustrated embodiments are for understanding purposes only and will not be repeated here.

[0094] This embodiment also provides a computer storage medium storing computer-executable instructions that can execute the two-factor authentication method for inducing content publication described in any of the above method embodiments. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk drive (HDD), or solid-state drive (SSD), etc.; the storage medium may also include combinations of the above types of memory.

[0095] Finally, it should be noted that the above preferred embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail through the above preferred embodiments, those skilled in the art should understand that various changes can be made to it in form and detail without departing from the scope defined by the claims of the present invention.

Claims

1. A two-factor authentication method for inducing content posting, characterized in that, include: Acquire pre-release processed data and pre-release raw data. The pre-release processed data represents the data stored in the database after the pre-release raw data has been processed according to the format requirements. Based on the pre-release processed data and the pre-release original data, content consistency is determined; If the content is consistent, perform anomaly detection on the pre-release processing data; If there are no abnormalities, the pre-release processing data will be sent to the inducing release device so that the inducing release device can display it; Images are collected and displayed on the device used for data acquisition and distribution. Based on the images displayed by the induced release device and the corresponding pre-released processed data, an emergency strategy is determined. The pre-release processing data is a combination of text and images. Before determining content consistency based on the pre-release processing data and the original pre-release data, the following is included: Perform image-text segmentation on the pre-released processed data to obtain the target text and target image; Based on a pre-trained target language model, text focus is extracted from the target text. Text focus includes the text body, action, and condition. Based on a pre-trained target detection model, visual focus anchors in target images are determined. Visual focus anchors include target anchors, action anchors, and numerical anchors. The subject in the target text and the corresponding target anchor point in the target image are input into a pre-trained target machine learning model to determine the vector similarity between the subject and the target anchor point in the semantic space. The action in the target text is compared with the action represented by the corresponding action anchor point in the target image, and the action matching degree is calculated. The conditions in the target text are parsed into logical expressions, and the values ​​represented by the corresponding numerical anchors in the target image are substituted into the corresponding logical expressions to determine the degree of numerical compliance. Based on vector similarity, action matching degree, and numerical consistency degree, determine whether to execute the content consistency determination step based on pre-release processed data and pre-release original data; The action matching degree is calculated by comparing the actions in the target text with the actions represented by the corresponding action anchor points in the target image; this includes: Based on the actions in the target text, determine whether it contains directional information; When directional information is included, the directional information is encoded into a first directional vector based on predefined encoding rules; Based on the image gradient algorithm, the direction of the corresponding action anchor point in the target image is determined; Based on predefined encoding rules, the direction of the action anchor point is encoded to obtain the second direction vector; Determine the dot product result based on the first direction vector and the second direction vector; Based on the dot product results, the action matching degree is determined.

2. The two-factor authentication method for inducing content publication according to claim 1, characterized in that, The pre-release processed data consists of images. Based on the pre-release processed data and the original pre-release data, content consistency is determined, including: Parse the pre-release raw data structure and extract the first visual elements, which include the position and size of the shape, the content and position of the text, the RGB values ​​of the color, and the element hierarchy; Based on image recognition algorithms, pre-released data is identified and second visual elements are extracted. These second visual elements include the position and size of shapes, the content and position of text, the RGB values ​​of colors, and the element hierarchy. The first visual element and the second visual element are quantitatively compared to obtain the consistency verification result. The quantitative comparison includes the existence of the element, positional deviation, size deviation, text matching degree, RGB, and element level matching degree.

3. The two-factor authentication method for inducing content publication according to claim 1, characterized in that, The pre-release processed data consists of images. Based on the pre-release processed data and the original pre-release data, content consistency is determined, including: Image recognition algorithms are used to identify images and obtain text information from them. The pre-released raw data is input into a pre-trained first intent recognition model to obtain the intent of the pre-released raw data. The text information is input into a pre-trained second intent recognition model to obtain the intent of the pre-published processing data; Match the intent of the pre-released raw data with the intent of the pre-released processed data to determine content consistency.

4. The two-factor authentication method for inducing content publication according to claim 1, characterized in that, Action anchors are arrow-shaped. Based on an image gradient algorithm, the direction of the corresponding action anchor in the target image is determined, including: Based on the detection coordinates of arrow-type action anchor points, a rectangular ROI region containing only arrows is cropped from the target image; The ROI region is converted into a single-channel grayscale image to obtain the first image; The contrast of the first image is enhanced by adaptive histogram equalization to obtain the second image; A thresholding algorithm is used to process the second image to obtain a binary image; The binary image is convolved using the Sobel operator to obtain the directional gradient matrix; Based on the pixel magnitude in the first image, determine the gradient magnitude and gradient direction angle of the pixel; Valid pixels are determined based on a pre-set gradient magnitude threshold and the gradient magnitude of the pixel. The gradient direction angles of the effective pixels are divided to obtain multiple directional partitions; Calculate the effective pixel gradient magnitude and the weighted sum of pixel values ​​within each directional interval; The direction interval containing the maximum weighted sum is taken as the dominant direction interval to obtain the direction of the action anchor point.

5. The two-factor authentication method for inducing content publication according to claim 4, characterized in that, The gradient direction angles of the effective pixels are divided into multiple direction partitions, including: Based on image segmentation algorithms, arrow regions are segmented from a binary image; Calculate the width and length of the arrow region; Select the maximum width value as the target width value, and select the maximum length value as the target length value; The number of directional partitions is determined based on the target length and target width values; Based on the number of directional partitions, the gradient direction angles of the effective pixels are divided to obtain multiple directional partitions.

6. A two-factor authentication device for inducing content posting, characterized in that, include: The data acquisition module is used to acquire pre-release processed data and pre-release raw data. The pre-release processed data represents the data stored in the database after the pre-release raw data has been processed according to the format requirements. The consistency determination module is used to determine content consistency based on pre-release processed data and pre-release original data; The anomaly detection module is used to perform anomaly detection on the pre-release processed data when the content is consistent. The sending module is used to send the pre-release processing data to the inducing release device when there are no abnormalities, so that the inducing release device can display it; The display module is used to collect images displayed by the inducement and release device; The emergency strategy determination module is used to determine the emergency strategy based on the image displayed by the inducing release device and the corresponding pre-release processing data; The pre-release processing data is a combination of text and images. Before determining content consistency based on the pre-release processing data and the original pre-release data, the following is included: Perform image-text segmentation on the pre-released processed data to obtain the target text and target image; Based on a pre-trained target language model, text focus is extracted from the target text. Text focus includes the text body, action, and condition. Based on a pre-trained target detection model, visual focus anchors in target images are determined. Visual focus anchors include target anchors, action anchors, and numerical anchors. The subject in the target text and the corresponding target anchor point in the target image are input into a pre-trained target machine learning model to determine the vector similarity between the subject and the target anchor point in the semantic space. The action in the target text is compared with the action represented by the corresponding action anchor point in the target image, and the action matching degree is calculated. The conditions in the target text are parsed into logical expressions, and the values ​​represented by the corresponding numerical anchors in the target image are substituted into the corresponding logical expressions to determine the degree of numerical compliance. Based on vector similarity, action matching degree, and numerical consistency degree, determine whether to execute the content consistency determination step based on pre-release processed data and pre-release original data; The action matching degree is calculated by comparing the actions in the target text with the actions represented by the corresponding action anchor points in the target image; this includes: Based on the actions in the target text, determine whether it contains directional information; When directional information is included, the directional information is encoded into a first directional vector based on predefined encoding rules; Based on the image gradient algorithm, the direction of the corresponding action anchor point in the target image is determined; Based on predefined encoding rules, the direction of the action anchor point is encoded to obtain the second direction vector; Determine the dot product result based on the first direction vector and the second direction vector; Based on the dot product results, the action matching degree is determined.

7. An electronic device, the device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor performs the steps of the two-factor authentication method for induced content publishing as described in any one of claims 1-5.

8. A computer storage medium storing computer instructions thereon, characterized in that, When executed by the processor, this instruction implements the steps of the two-factor authentication method for induced content publishing as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Airport terminal display screen content auditing method, device and equipment and medium

    CN121262387A