Bid evaluation site multi-modal behavior compliance monitoring method and system
By simultaneously collecting sound, image, and location information at the bid evaluation site, and utilizing multimodal analysis technology to monitor violations in the bid evaluation process in real time and automatically intervene, the problem of traditional bid evaluation monitoring systems being unable to detect hidden violations has been solved, achieving real-time compliance control and fairness assurance of the bid evaluation process.
Patent Information
- Application Number
- CN202511128129.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-11-07
AI Technical Summary
Traditional bid evaluation monitoring systems struggle to detect hidden violations in the bid evaluation process in real time, and lack a unified behavior compliance scoring and real-time intervention mechanism, leading to the risk of opaque operations in the bid evaluation process.
The system uses array microphones and high-definition cameras to simultaneously collect sound and image data, and combines Bluetooth Low Energy or UWB tags to collect personnel location in real time. It uses streaming speech recognition, posture estimation, and object recognition algorithms to judge language and body behavior, calculates on-site compliance scores through a rule engine, and intervenes in real time and preserves evidence when violations occur.
It achieves comprehensive, real-time, and traceable compliance control over the bidding process, significantly reducing the risk of opaque operations and unauthorized communication, and ensuring the fairness and impartiality of the bidding process.
Smart Images

Figure CN120913544A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence monitoring technology, in particular to a method and system for monitoring multi-modal behavior compliance in bid evaluation site. BACKGROUND
[0002] In government procurement and large-scale enterprise bidding activities, the on-site behavior of bid evaluation experts and staff directly affects the fairness and legality of bid evaluation results.
[0003] Traditional prevention and control methods mainly rely on manual supervision and video recording for post-checking, which is difficult to discover hidden irregular actions such as private communication, suggestive language, hidden exchange of documents or mobile phone photography in time. Even if discovered, it is often difficult to obtain evidence due to the lack of synchronous audio and video evidence. At the same time, the existing monitoring system usually manages voice, video and personnel positioning as independent modules, which cannot form a unified behavior compliance scoring and real-time intervention mechanism, resulting in the risk of dark box operation in the bid evaluation process.
[0004] Therefore, a comprehensive monitoring technology is needed that can capture multi-modal irregular behavior in real time in the bid evaluation site and automatically take intervention measures to improve the transparency and efficiency of bid evaluation. SUMMARY
[0005] The purpose of the present application is to provide a method and system for monitoring multi-modal behavior compliance in bid evaluation site to solve the problems raised in the background technology.
[0006] To achieve the above purpose, the present application provides the following technical solution: a method for monitoring multi-modal behavior compliance in bid evaluation site, comprising the following steps:
[0007] a) deploying array microphones and high-definition cameras inside the bid evaluation room to synchronously collect voice and image data of bid evaluation experts and staff, and simultaneously collecting personnel position information through Bluetooth Low Energy or Ultra-Wideband tags in real time;
[0008] b) inputting the collected voice data into a streaming speech recognition model to transcribe it into text, and then detecting the bidder's name, bid amount and preset sensitive keywords by a pre-trained language model, and outputting a language risk score; wherein the language model automatically updates the tender project name and industry jargon through few-shot prompting, without the need for manual maintenance of a fixed sensitive word library;
[0009] c) input the camera image data into pose estimation, object recognition and eye tracking algorithms to determine whether there is body contact, article transfer, mobile phone use or long gaze at others' screens abnormal behavior, and output the action risk score; wherein the action recognition uses pose skeleton point distance and angle to determine body contact, and combines object detection results to confirm file, mobile storage or mobile phone transfer behavior; the eye tracking algorithm calculates the line of sight vector according to the head pose and pupil direction, and marks a suspicious gaze event when the line of sight stays in others' screen or score table area for more than five seconds;
[0010] d) compare the real-time coordinates of the personnel with the seat reference points, and if the time away from the seat exceeds the preset time and no clock-in is performed, generate a risk score of leaving the seat; the risk score of leaving the seat is compensated by the multi-target tracking results of the camera when the Bluetooth low-power tag signal is temporarily lost, to avoid false positives;
[0011] e) input the language risk score, action risk score and leave seat risk score into the rule engine to calculate the on-site compliance score R according to the configured weights; the rule engine uses a gated recurrent unit to smooth and de-duplicate repeated events within the same window, and amplifies the corresponding risk weight when language violations and article transfer occur at the same time;
[0012] f) send a yellow reminder when the score R is below the first threshold, and automatically lock the evaluation terminal of the violating personnel and push an alarm information to the supervision department when the score R is below the second threshold; when the evaluation terminal of the violating personnel is locked, the system simultaneously intercepts the violating frame image and caches the audio segments of the previous and next thirty seconds as complete evidence; the first threshold and the second threshold can be dynamically adjusted through the background management interface, and the system can take effect without restarting after the threshold is updated;
[0013] g) write all event records, compliance score curves and audio-video segment indexes into a chain hash log, and support one-key export of audit reports; the chain hash log uses a hash chain structure, and the hash value of each log record is calculated from the previous hash value and the current record content, to ensure that the log is tamper-proof.
[0014] Preferably, the chain hash log is desensitized and used as a training sample to periodically fine-tune the language model and action recognition model, to improve the accuracy of violation behavior recognition.
[0015] Preferably, in step b, the pre-trained language model automatically updates the tender project name and industry jargon through few-shot prompting, without the need for manual maintenance of a fixed sensitive word library, to adapt to the needs of different tender projects.
[0016] Preferably, in step c, the action recognition uses pose skeleton point distance and angle to judge limb contact, and combines object detection results to confirm file, mobile storage or mobile phone transfer behavior; the eye tracking algorithm calculates the line of sight vector according to the head pose and pupil direction, and marks a suspicious gaze event when the line of sight stays in the other person's screen or score table area for five consecutive seconds, to accurately judge abnormal behavior.
[0017] Preferably, in step d, the off-site risk score is compensated by the multi-target tracking result of the camera when the Bluetooth low-power tag signal is temporarily lost, to avoid false positives caused by signal loss and improve the accuracy of off-site risk judgment.
[0018] A system for a multi-modal behavior compliance monitoring method in a bid evaluation site, comprising:
[0019] A data acquisition module for synchronously acquiring sound and image data of bid evaluation experts and staff in the bid evaluation room by deploying array microphones and high-definition cameras, and acquiring personnel position information in real time by Bluetooth low-power or ultra-wideband tags;
[0020] A language risk analysis module for inputting the sound data collected by the data acquisition module into a streaming speech recognition model to transcribe the sound data into text, and then detecting the tenderer name, bid number and preset sensitive keywords by a pre-trained language model, and outputting a language risk score;
[0021] An action risk analysis module for inputting the image data collected by the camera in the data acquisition module into pose estimation, object recognition and eye tracking algorithms to judge whether there are abnormal behaviors such as limb contact, article transfer, mobile phone use or long-time gaze at other people's screens, and outputting an action risk score; wherein the action recognition uses pose skeleton point distance and angle to judge limb contact, and combines object detection results to confirm file, mobile storage or mobile phone transfer behavior; the eye tracking algorithm calculates the line of sight vector according to the head pose and pupil direction, and marks a suspicious gaze event when the line of sight stays in the other person's screen or score table area for five consecutive seconds;
[0022] An off-site risk analysis module for comparing the real-time coordinates of the personnel with the seat reference points, and generating an off-site risk score if the off-site time exceeds the preset time length and no clock-in is performed; the off-site risk score is compensated by the multi-target tracking result of the camera when the Bluetooth low-power tag signal is temporarily lost, to avoid false positives;
[0023] A score calculation module for inputting the language risk score output by the language risk analysis module, the action risk score output by the action risk analysis module and the off-site risk score output by the off-site risk analysis module into a rule engine to calculate an on-site compliance score R according to the configured weights; the rule engine uses a gated recurrent unit to smooth and de-duplicate repeated events in the same window, and amplifies the corresponding risk weight when a language violation and article transfer occur at the same time.
[0024] The early warning processing module sends a yellow reminder when the score R calculated by the score calculation module is lower than a first threshold value, and automatically locks the evaluation terminal of the irregular personnel and pushes an alarm information to the supervisory department when the score R is lower than a second threshold value; when the evaluation terminal of the irregular personnel is locked, the system simultaneously intercepts the irregular frame image and caches the audio segments of thirty seconds before and after, and saves them as complete evidence; the first threshold value and the second threshold value can be dynamically adjusted through a background management interface, and the system can take effect without restarting after the threshold values are updated;
[0025] The log recording module writes all event records, compliance score curves and audio and video segment indexes into a chain hash log, and supports one-key export of an audit report; the chain hash log adopts a hash chain structure, and the hash value of each log record is calculated by the hash value of the previous record and the current record content, so that the log is ensured to be tamper-proof.
[0026] Preferably, the language model used in the language risk analysis module is automatically updated with few sample prompts to adapt to the needs of different bidding projects and improve the flexibility of detection without manually maintaining a fixed sensitive word library.
[0027] Preferably, the action risk analysis module accurately judges the abnormal behavior of the evaluation site personnel by comprehensively considering multi-dimensional information such as the distance and angle of the posture skeleton point, the object detection result and the line of sight vector, thereby improving the accuracy and reliability of the action risk judgment.
[0028] Preferably, the off-site risk analysis module effectively solves the false alarm problem caused by signal loss through the cooperative work of the Bluetooth low-power tag and the multi-target tracking of the camera, thereby ensuring the accuracy of the off-site risk judgment.
[0029] Preferably, the system further comprises a model fine-tuning module, which is used for desensitizing the chain hash log in the log recording module as a training sample, and periodically fine-tuning the language model in the language risk analysis module and the action recognition model in the action risk analysis module to improve the accuracy of the irregular behavior recognition.
[0030] Compared with the prior art, the system has the following beneficial effects:
[0031] This invention proposes a multimodal behavior compliance monitoring method and system for bid evaluation sites. By simultaneously collecting on-site sound, images, and personnel location information, it monitors in real-time the language, body language, object transfer, and electronic device usage behaviors of bid evaluation experts and staff during the bid evaluation process. The system first uses an array microphone and a speech recognition model to transcribe on-site speech into text, and then automatically detects the presence of bidder names, bid numbers, or other sensitive keywords using a large language model. Simultaneously, it employs high-definition cameras, posture estimation, and object recognition algorithms to capture abnormal actions such as whispering among judges, physical contact, exchange of documents or storage media, and unauthorized use of mobile phones for photography. The system also uses a positioning module to detect the duration of unauthorized absences by judges and eye-tracking technology to identify judges' prolonged staring at other people's screens. Data on voice violations, abnormal actions, excessive absences, and equipment violations are input into a rule engine to calculate on-site compliance scores and generate real-time alerts. When the score falls below a set threshold, the system automatically suspends the bid evaluation account and sends an alarm to the supervisory department. Monitoring results are saved in the form of event screenshots, voice-text chains, and logs, supporting later auditing and evidence collection. This invention can provide comprehensive, real-time, and traceable compliance control at the bid evaluation site, significantly reducing the risk of opaque operations and unauthorized communication, and ensuring the fairness and impartiality of the bid evaluation process. Attached Figure Description
[0032] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation
[0033] To make the objectives, technical solutions, and advantages of the present invention clear and complete, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only some, not all, embodiments of the present invention, and are merely illustrative of the embodiments of the present invention. They are not intended to limit the embodiments of the present invention. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0034] Example 1: This invention provides a technical solution: a method for monitoring multimodal behavior compliance at a bid evaluation site, comprising the following steps:
[0035] Step 1: Audio Acquisition and Semantic Analysis
[0036] Before the evaluation process begins, the array microphones are activated and enter continuous monitoring mode. The first processing stage sends the raw sound waves to the beamforming module, which automatically distinguishes between near and far sound sources and eliminates ambient noise. This not only ensures the clarity of subsequent identification but also lays the foundation for sound source localization.
[0037] The noise-reduced speech stream then enters the streaming speech recognition model, which uses a chunked inference approach to output text snippets and confidence scores every 200 milliseconds. The text snippets immediately enter the semantic filter, which uses a large language model to identify bidder names, bid amounts, or suggestive passphrases using few-shot prompts and calculate sensitive word intensity. If two high-intensity sensitive words appear consecutively within the same window, the system temporarily stores the corresponding audio and pushes the event to the fusion layer.
[0038] Meanwhile, acoustic features such as volume, speech rate, and pitch are extracted into statistical vectors that are fed into an emotion classification network. The network outputs three labels - calm, tense, and excited - along with a confidence score. If a judge or staff member's emotional state rapidly switches from calm to excited or whispers for more than 3 seconds, the system generates a "abnormal voice emotion" event for further decision-making.
[0039] Step 2: Video Capture and Behavior Recognition
[0040] Two 4K spherical cameras cover the evaluation area and the question-answering area, with full-frame input target detection and multi-target tracking models. The model assigns a unique ID to all present personnel and outputs head-shoulder boxes and skeleton coordinates, ensuring continuous tracking even with partial occlusion.
[0041] On the skeleton data, the pose estimation algorithm captures limb contact or object transfer using hand-to-torso distance and relative angle changes. If two people are detected with hands within 30 cm and accompanied by the appearance of paper, mobile storage, or mobile phones, the system immediately intercepts 10 high-resolution images and sends a "object transfer" event.
[0042] The video stream is also branched to an eye tracking sub-model. This model calculates the gaze vector from head pose and pupil direction. When a judge's gaze falls on another judge's screen or score sheet location for more than 5 seconds, the system labels it as "suspicious gaze" and records the corresponding coordinates and time interval for later verification of plagiarism or hints.
[0043] Step 3: Location Awareness and Leave Monitoring
[0044] The BLE tag worn by each judge broadcasts its ID at 1-second intervals, and the positioning base station calculates the triangulation coordinates in real time and writes them to the location buffer. The system maintains a seat reference table and compares the real-time coordinates to the distance from the seat center point. If the distance exceeds 1 meter and lasts for 60 seconds without terminal check-in, the system generates an "unauthorized leave" event.
[0045] To improve positioning robustness, the system cross-verifies BLE data and video skeleton tracking results. When the tag signal is temporarily lost, the video tracking can still output the person's trajectory and provide a "seat replacement" result, avoiding false alarms.
[0046] The location data is also interpreted in conjunction with the video events. When the off-site person returns and is accompanied by an item delivery action, the system merges the off-site event with the delivery event, resulting in a higher-weighted composite risk event.
[0047] Step 4: Event fusion and risk scoring
[0048] The events generated by the three channels of audio, video, and location are all in a unified format, including event type, timestamp, object ID, and confidence. The messages enter a gating loop unit for time window aggregation, first smoothing out similar repeated events in a short period of time to prevent the same behavior from being counted multiple times.
[0049] Subsequently, the events are sent to the rule engine. The engine reads a configurable weight table: for example, language violations account for 40%, body or item delivery 30%, electronic devices and off-site 15%. The engine weights and accumulates the event confidence according to the weights to obtain a real-time compliance score. If a keyword trigger and an item delivery occur simultaneously within a 30-second window, the rule engine amplifies the risk score by a cross coefficient to reflect the event correlation.
[0050] The compliance score is displayed in a curve form on the supervisor's seat panel and written to Redis cache in the background for easy query by other modules. Each drop in the curve is annotated with a reason text, and the supervisor can view the sound clip or violation screenshot with one click, enhancing decision transparency.
[0051] Step 5: Active intervention and alarm execution
[0052] When the compliance score falls below the threshold of 80 points, the system switches the indicator light from green to yellow and pushes a "Please pay attention" reminder to the supervisor's seat through WebSocket, while a flashing bar pops up at the top of the evaluation system interface. If the score continues to drop below 60 points, the system enters a red alarm state: the background immediately locks the violation person's score interface, preventing further input, and automatically archives the violation frame and voice text.
[0053] After the red alarm is triggered, the intervention module sends a text message and email notification to the evaluation management end at the same time, containing the violation type, time, and person ID. The supervisor needs to input an unlock password or complete on-site review to resume the evaluation progress, otherwise the system remains locked and continues to record video.
[0054] If the on-site detects the use of a violation electronic device, the system can drive the camera to automatically zoom to the violation device screen and start a dedicated screen recording channel, so that the screen operation process is completely retained. All intervention actions and operation instructions are recorded in the transaction log, providing a basis for post-responsibility determination.
[0055] Step 6: Evidence archiving and model iteration
[0056] When the evaluation is over, the log processing module writes all events, compliance curves, and audio-video segment indexes into the immutable log table in chronological order, and concatenates the entries in a hash chain to generate the final chain head value, which is saved to the read-only repository. Any later tampering will cause the chain head to be inconsistent, providing technical-level anti-fraud protection.
[0057] The system also sanitizes the event data, removing bidder information and sensitive amounts, and only retaining behavior tags and feature vectors. These sanitized samples are sent to the training repository and used periodically to update the language model sensitive word set, the video model action threshold, and the emotion classifier weight. The update process uses an A / B grayscale strategy, with the new model running in shadow mode, and only switching to the online version when the accuracy improves by more than 5% without increasing the false positive rate.
[0058] Through continuous data closed-loop and model adaptation, the system not only ensures real-time and fair evaluation, but also can deal with new types of violations, maintaining high accuracy and low interference rate in the long term.
[0059] In embodiment two, based on embodiment one, a system for multi-modal behavior compliance monitoring in the evaluation site is proposed, comprising:
[0060] The audio acquisition and semantic analysis layer, the video acquisition and behavior recognition layer, the location perception layer, the data fusion and risk scoring layer, the intervention and alarm layer, and the evidence archiving and auditing layer. Each layer can be independently deployed or combined into a complete behavior compliance management platform through a message bus.
[0061] 1. Audio acquisition and semantic analysis layer
[0062] This layer arranges array microphones and front-end edge boxes in the evaluation room. The microphones first separate the far-end and near-end sound sources using beamforming technology, and then convert the conversation into text in real time through a streaming speech recognition model. After the transcription is completed, the text immediately enters the large language model semantic filter. The model matches high-risk keywords such as bidder names, bid amounts, and suggestive instructions in a Few-Shot Prompt manner and outputs a confidence score.
[0063] At the same time, acoustic features (volume, speech rate, pitch) are sent to the emotion detection network to determine whether the speaker is in an agitated, angry, or deliberately low-volume state. If continuous whispering or emotional changes are detected, the system will generate an audio event tag within milliseconds and push it to the data fusion layer for subsequent decision-making.
[0064] 2. Video acquisition and behavior recognition layer
[0065] The video layer uses dual 4K ball cameras to cover the evaluation area and the Q&A area. Image data is first processed by a multi-target tracking algorithm to generate personnel trajectories, and then fed into a pose estimation network to extract skeleton points. Based on the distance and angle of the skeleton, the system can determine within two seconds whether there is body contact, document exchange, or mobile phone shooting, and distinguish between ordinary body swings and suspicious transfer behaviors.
[0066] To prevent staff from bypassing the camera and blocking, the system adds an object re-identification and obstruction recovery mechanism to the model: when the key body nodes disappear or are blocked, the algorithm automatically switches to the next camera view and continues tracking. At the same time, the eye tracking submodule analyzes the gaze direction and dwell time of the judges. If they stare at another judge's screen or score sheet for a long time, it is marked as a suspicious gaze event.
[0067] 3. Position-aware layer
[0068] Each judge wears a BLE or UWB tag, and the positioning base stations are arranged with a spacing of less than 0.3 meters to ensure accurate positioning. The tag coordinates are continuously compared with the reference points of the judges' seats: if the system determines that the judge has left their seat for more than 60 seconds without performing a check-in at the system terminal, the position layer generates an off-site abnormal event.
[0069] In addition, the position data and video skeleton trajectories are cross-verified to avoid false positives due to temporary loss of tag signals. After fusion layer confirmation, the off-site event is written into the real-time risk score and highlighted on the large screen panel to alert supervisors.
[0070] 4. Data fusion and risk scoring layer
[0071] The fusion layer receives event streams from audio, video, and position using a Kafka message bus. Each event has a timestamp and a confidence level. The events first pass through a gated recurrent unit (GRU) to smooth fluctuations, and then enter the rule engine. The rule engine maintains a configurable weight table: for example, language sensitivity accounts for 40%, body contact accounts for 30%, and off-site and electronic device violations each account for 15%.
[0072] The engine accumulates scores according to weights to obtain an on-site "compliance score". If the audio layer and video layer report keyword triggers and document transfers within a 30-second window, the system applies a multiplicative weighting strategy to increase the corresponding risk score, avoiding single-point false positives. In addition, the compliance score is dynamically decayed over time to ensure that a short action does not permanently lock the risk.
[0073] 5. Intervention and alarm layer
[0074] When the compliance score is below the preset threshold of 80 points, the system sends a yellow warning to the supervision seat through WebSocket; below 60 points, it enters a red alarm state, automatically locks the screen of the evaluator terminal in the background of the bid evaluation system, suspends the scoring operation, and pops up a window with screenshots and voice texts of the violation frame to the supervision group for confirmation. The intervention instructions are recorded in the background for the complete life cycle, and can be released or upgraded according to the instructions of the supervision group.
[0075] For electronic device violations or whispering events, the system can automatically switch the camera to the violator and start high-frame-rate recording, while buffering the current audio and video segments for 30 seconds before and after to ensure complete evidence.
[0076] 6. Evidence storage and audit layer
[0077] All event data, compliance score change curves, intervention logs, and audio and video segments are written into a log database that only increases and does not decrease, and chain hash archives are used to ensure that any later modification will cause a hash chain break. The system supports one-key export of audit packages, including CSV event lists and embedded screenshots, audio files, which can be directly submitted to the discipline inspection or judicial departments.
[0078] After each bid evaluation is completed, the system desensitizes the logs and sends them to the training warehouse to update the sensitive word set of the language model and the posture threshold of the video model, realizing continuous self-learning. This not only ensures the safety and compliance of the bid evaluation data, but also improves the model's adaptability to new types of violations.
[0079] Although embodiments of the present application have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and alterations can be made without departing from the principles and spirit of the present application, and the scope of the present application is defined by the appended claims and their equivalents.
Claims
1. A method for bid evaluation site multi-modal behavior compliance monitoring, characterized in that: Comprising the following steps: a) Deploying an array microphone inside the bid evaluation room to synchronously collect the sound and image data of the bid evaluation experts and staff with a high-definition camera, and collecting the personnel position information in real time through Bluetooth low power or ultra-wideband tags; b) Inputting the collected sound data into a streaming speech recognition model to transcribe it into text, and then detecting the bidder's name, bid number, and preset sensitive keywords by a pre-trained language model, and outputting a language risk score; wherein the language model automatically updates the bid project name and industry jargon through few-shot prompting, without the need for manual maintenance of a fixed sensitive word library; c) Inputting the camera image data into a pose estimation, object recognition, and eye tracking algorithm to determine whether there are abnormal behaviors such as body contact, object transfer, mobile phone use, or long-time gaze at others' screens, and outputting a motion risk score; wherein the motion recognition uses pose skeleton point distance and angle to judge body contact, and combines with object detection results to confirm file, mobile storage, or mobile phone transfer behavior; the eye tracking algorithm calculates the gaze vector according to the head pose and pupil direction, and marks a suspicious gaze event when the gaze stays in others' screen or scoring table area for more than five seconds; d) Comparing the real-time coordinates of the personnel with their seat reference points, and if the time away from the seat exceeds the preset time and no clock-in is performed, a time-away risk score is generated; the time-away risk score is compensated by the multi-target tracking results of the camera when the Bluetooth low power tag signal is temporarily lost, to avoid false positives; e) Inputting the language risk score, motion risk score, and time-away risk score into a rule engine to calculate the on-site compliance score R according to the configured weights; the rule engine uses a gated recurrent unit to smooth and de-duplicate repeated events within the same window, and amplifies the corresponding risk weight when there is a language violation and object transfer at the same time; f) Sending a yellow alert when the score R is below a first threshold, and automatically locking the bid evaluation terminal of the violating personnel and pushing an alarm information to the supervision department when the score R is below a second threshold; when the bid evaluation terminal of the violating personnel is locked, the system simultaneously intercepts the violating frame image and caches the audio segments of the previous and next thirty seconds as complete evidence; the first threshold and the second threshold can be dynamically adjusted through the background management interface, and the system can take effect without restarting after the threshold is updated; g) Writing all event records, compliance score curves, and audio-video segment indexes into a chain hash log, and supporting one-key export of audit reports; the chain hash log uses a hash chain structure, and the hash value of each log record is calculated from the previous hash value and the current record content, ensuring that the log is tamper-proof.
2. The method for monitoring multi-modal behavior compliance in a bid evaluation site according to claim 1, wherein: Further comprising periodically fine-tuning the language model and the motion recognition model using the desensitized chain hash log as training samples to improve the accuracy of the violation behavior recognition.
3. The method for monitoring multi-modal behavior compliance in a bid evaluation site according to claim 2, wherein: In step b, the pre-trained language model automatically updates the bid project name and industry jargon through few-shot prompting, without the need for manual maintenance of a fixed sensitive word library, to adapt to the needs of different bid projects.
4. The method for monitoring multi-modal behavior compliance in a bid evaluation site according to claim 3, wherein: In step c, action recognition uses pose skeleton point distance and angle to determine limb contact, and combines object detection results to confirm document, mobile storage or mobile phone transfer behavior; eye tracking algorithm calculates gaze vector according to head pose and pupil direction, and marks suspicious gaze events when gaze stays in other people's screen or score table area for more than five seconds, to accurately determine abnormal behavior.
5. The method for monitoring multi-modal behavior compliance in a bid site as claimed in claim 4, wherein: In step d, when the Bluetooth low energy tag signal is temporarily lost, the off-site risk score is compensated by the multi-target tracking result of the camera, to avoid false positives caused by signal loss and improve the accuracy of off-site risk judgment.
6. A system for bid site multi-modal behavioral compliance monitoring method as claimed in claim 5, wherein: It comprises: A data acquisition module for synchronously collecting voice and image data of bid evaluation experts and staff inside the bid evaluation room through array microphones and high-definition cameras, and collecting real-time personnel position information through Bluetooth low energy or ultra-wideband tags; A language risk analysis module for inputting voice data collected by the data acquisition module into a streaming speech recognition model to transcribe the voice data into text, and detecting bidder names, bid amounts and preset sensitive keywords through a pre-trained language model, and outputting a language risk score; An action risk analysis module for inputting image data collected by the camera in the data acquisition module into pose estimation, object recognition and eye tracking algorithms to determine whether there are abnormal behaviors such as limb contact, object transfer, mobile phone use or long-time gaze at other people's screens, and outputting an action risk score; wherein action recognition uses pose skeleton point distance and angle to determine limb contact, and combines object detection results to confirm document, mobile storage or mobile phone transfer behavior; eye tracking algorithm calculates gaze vector according to head pose and pupil direction, and marks suspicious gaze events when gaze stays in other people's screen or score table area for more than five seconds; An off-site risk analysis module for comparing real-time coordinates of personnel with their seat reference points, and generating an off-site risk score if the off-site time exceeds the preset time and no clock-in is performed; when the Bluetooth low energy tag signal is temporarily lost, the off-site risk score is compensated by the multi-target tracking result of the camera to avoid false positives; A score calculation module for inputting the language risk score output by the language risk analysis module, the action risk score output by the action risk analysis module and the off-site risk score output by the off-site risk analysis module into a rule engine to calculate an on-site compliance score R according to the configured weights; the rule engine uses a gated recurrent unit to smooth and de-duplicate repeated events in the same window, and amplifies the corresponding risk weight when language violations and object transfer occur at the same time; An early warning processing module for sending a yellow reminder when the score R calculated by the score calculation module is lower than a first threshold, and automatically locking the bid evaluation terminal of the violating personnel and pushing alarm information to the supervision department when the score R is lower than a second threshold; when the bid evaluation terminal of the violating personnel is locked, the system simultaneously intercepts the violating frame image and caches audio clips of thirty seconds before and after, as complete evidence; the first threshold and the second threshold can be dynamically adjusted through a background management interface, and the system can take effect without restarting after the threshold is updated; The log recording module writes all event records, compliance score curves and audio-video clip indexes into a chain hash log, and supports one-key export of audit reports; the chain hash log adopts a hash chain structure, a hash value of each log record is calculated from a previous hash value and current record content, and the log is ensured to be tamper-proof.
7. The system of claim 6, wherein: The language model used in the language risk analysis module automatically updates the tender project name and industry jargon through few-shot prompting, without the need for manual maintenance of a fixed sensitive word library, to adapt to the needs of different tender projects and improve the flexibility of detection.
8. The system of claim 7, wherein: The action risk analysis module accurately judges the abnormal behavior of the personnel on the bid evaluation site by comprehensively considering multi-dimensional information such as the distance and angle of the posture skeleton points, the object detection result and the line of sight vector, thereby improving the accuracy and reliability of action risk judgment.
9. The system of claim 8, wherein: The off-site risk analysis module effectively solves the false alarm problem caused by signal loss through the cooperative work of the Bluetooth low-power tag and the multi-target tracking of the camera, thereby ensuring the accuracy of off-site risk judgment.
10. The system of claim 9, wherein: The model fine-tuning module is further included for desensitizing the chain hash log in the log recording module as a training sample, and periodically fine-tuning the language model in the language risk analysis module and the action recognition model in the action risk analysis module, to improve the accuracy of the violation behavior recognition.
Citation Information
Cited By
Bid evaluation base personnel abnormal behavior monitoring method based on face recognition
CN121884462A
Behavior analysis system and analysis method
CN122290220A