Behavior early warning and digital witness method and device for bid evaluation site
By collecting multimodal data at the bid evaluation site and aligning it with timestamps, and combining it with a deep learning model for risk assessment, the problem of insufficient intelligence in existing technologies is solved, and intelligent and reliable recording of the bid evaluation process is achieved.
Patent Information
- Application Number
- CN202511019346.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-11-07
AI Technical Summary
The existing bidding evaluation and supervision technology is not intelligent enough and lacks a reliable recording mechanism for the whole process, which leads to inaccurate capture of behavior during the bidding evaluation process and wastes human resources.
By acquiring multimodal data (video, audio, operation logs, and environmental parameters) from the bidding evaluation site, and using timestamps for data fusion and alignment, a multimodal risk feature vector is generated. This vector is then combined with a long short-term memory network and an isolated forest model for risk assessment, enabling behavioral early warning and digital witnessing.
It has improved the intelligence level of on-site behavior early warning, ensured the temporal accuracy and semantic integrity of risk analysis, and achieved accurate monitoring and reliable recording of the bidding process.
Smart Images

Figure CN120911952A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence, and in particular to a method and device for behavior early warning and digital witnessing in a bid evaluation site. BACKGROUND
[0002] The existing bid evaluation supervision technology has the problems of insufficient intelligence level and lack of a full-process credible record mechanism. In the current bid evaluation supervision process, a single video monitoring device or on-site staff patrol method is usually relied on, which cannot effectively capture other information involved in the bid evaluation process, resulting in inaccurate bid evaluation site behavior capture results and consumption of human resources. Therefore, there is an urgent need for a more intelligent and more accurate intelligent bid evaluation supervision method. SUMMARY
[0003] The present application provides a method and device for behavior early warning and digital witnessing in a bid evaluation site, which solves the technical problem of insufficient intelligence level in the prior art.
[0004] To achieve the above-mentioned purpose, the present application adopts the following technical solutions:
[0005] In a first aspect, a method for behavior early warning and digital witnessing in a bid evaluation site is provided, comprising: acquiring multi-modal data with a time stamp in a bid evaluation site, the multi-modal data including: video data, audio data, operation log data, and environmental parameters. Extracting feature vectors of each type of multi-modal data, and performing data fusion and alignment based on the time stamp to obtain a multi-modal risk feature vector. Inputting the multi-modal risk feature vector into a risk assessment model to determine the risk level of the bid evaluation site. Based on the risk level of the bid evaluation site, risk behavior early warning and evidence storage are performed.
[0006] In combination with the first aspect described above, in a possible implementation manner, the multi-modal risk feature vector includes behavior time sequence features and static features; the extracting of feature vectors of each type of multi-modal data and the performing of data fusion and alignment based on the time stamp to obtain the multi-modal risk feature vector include: extracting behavior features of the video data, the audio data, and the operation log data, the behavior features including: video behavior features, voice behavior features, and operation behavior features; extracting static features of the operation log data and the environmental parameters; and the behavior features and the static features are time-aligned through a sliding window manner on a time axis to obtain the multi-modal risk feature vector.
[0007] In a possible implementation manner of the first aspect, the behavior features are extracted based on the video data, the audio data, and the operation log data, including: extracting a person contour in the video data by using an image recognition algorithm, and taking the extracted person mask feature and the person position coordinates as the video behavior features; extracting acoustic features in the audio data by using STFT and MFCC, and taking the acoustic features as the speech behavior features; cleaning the operation log data to output target log data, and taking the target log data as the operation behavior features.
[0008] In a possible implementation manner of the first aspect, the behavior features and the static features are time-aligned in a time axis by using a sliding window, including: initializing the time axis; in each time window, traversing time stamps of all the behavior features to find a time stamp closest to the time window; if the matching is successful, taking the matched behavior features corresponding to the time stamp as a group of multi-modal risk feature vectors together with the static features; if the matching fails, skipping the current window and entering the next window to repeat the traversal and matching process.
[0009] In a possible implementation manner of the first aspect, the multi-modal risk feature vectors are input into a risk assessment model to determine a risk level of the bid evaluation site, including: sorting the behavior features in the multi-modal risk feature vectors based on the time sequence between the groups of multi-modal risk feature vectors to obtain behavior time sequence features; inputting the behavior time sequence features into a long short-term memory network model to output a dynamic risk probability score; inputting the static features into an isolation forest model to output a static anomaly probability score; weighting the dynamic risk probability score and the static anomaly probability score according to importance to determine a risk probability score, and outputting the risk level of the bid evaluation site.
[0010] In a possible implementation manner of the first aspect, the dynamic risk probability score P norm satisfies the following formula:
[0011]
[0012] wherein, P lstm represents the dynamic risk probability score output by the long short-term memory network for the current window, μ L represents a historical average risk probability score output by the long short-term memory network, and σ L represents a historical standard deviation output by the long short-term memory network.
[0013] In a possible implementation manner of the first aspect, the risk probability score R score satisfies the following formula:
[0014] R score = α·P norm + β·Snorm
[0015] wherein, a is a time sequence behavior weight, b is a static feature weight, P norm is a normalized time sequence risk probability, S norm is a normalized static anomaly probability
[0016] In a possible implementation manner of the first aspect, based on the risk level of the bid site, risk behavior early warning and evidence storage are performed, including: based on the risk level, determining a corresponding level of early warning response; when the risk level is greater than or equal to a preset level threshold, audio data, video data, and operation log data in a first time period are intercepted; based on the audio data, the video data, and the operation log data in the first time period, an evidence package is stored, and the evidence package includes: a timestamp, a risk level, a confidence, audio data, video data, and operation log data.
[0017] In a possible implementation manner of the first aspect, after the multi-modal risk feature vector is input into the risk assessment model to determine the risk level of the bid site, the method further includes: in a case where the risk level of the bid site confirmed by the human being is consistent with the actual risk level of the bid site, taking the multi-modal risk feature vector used to determine the risk level of the bid site as training data to train and optimize iterations of the risk assessment model.
[0018] The second aspect provides an electronic device, including: a communication unit and a processing unit; the communication unit is configured to acquire multi-modal data with a timestamp of a bid site, the multi-modal data including: video data, audio data, operation log data, and environmental parameters; the processing unit is configured to extract a feature vector of each type of multi-modal data respectively, and perform data fusion and alignment based on the timestamp to obtain a multi-modal risk feature vector; the processing unit is further configured to input the multi-modal risk feature vector into a risk assessment model to determine a risk level of the bid site; and based on the risk level of the bid site, risk behavior early warning is performed.
[0019] The third aspect provides an electronic device, including: a processor and a storage medium; the storage medium includes instructions, and the processor is configured to execute the instructions to implement the method described in the first aspect and any possible implementation manner of the first aspect. The electronic device can be an electronic device, or a chip in the electronic device.
[0020] In a fourth aspect, the application provides a behavior early warning and digital witnessing system for a bid evaluation site, comprising a data acquisition device and an electronic device. The data acquisition device is configured to acquire multi-modal data with timestamps from the bid evaluation site, the multi-modal data comprising video data, audio data, operation log data, and environmental parameters. The electronic device is configured to extract feature vectors of each type of multi-modal data, and perform data fusion and alignment based on the timestamps to obtain a multi-modal risk feature vector. The electronic device is further configured to input the multi-modal risk feature vector into a risk assessment model to determine a risk level of the bid evaluation site. Based on the risk level of the bid evaluation site, a risk behavior early warning is performed.
[0021] In a fifth aspect, the application provides a computer-readable storage medium having instructions stored therein, which, when executed on an electronic device, cause the electronic device to perform the method described in the first aspect and any possible implementation manner of the first aspect.
[0022] In a sixth aspect, the application provides a computer program product comprising instructions, which, when executed on an electronic device, cause the electronic device to perform the method described in the first aspect and any possible implementation manner of the first aspect.
[0023] To address the technical problem of insufficient intelligence level in the prior art, the application provides a bid evaluation behavior early warning and witnessing method integrating multi-modal perception, intelligent behavior analysis, and digital witnessing. Compared with the conventional method relying on manual patrol or single video monitoring, the technical solution provided by the application can simultaneously acquire multi-modal data such as audio, video, operation log, and environmental state of the bid evaluation site, realize spatio-temporal alignment of the data through a sliding window mechanism, and generate behavior time sequence features with complete semantics, thereby solving the technical problem of insufficient intelligence level in the prior art.
[0024] It should be understood that the description of technical features, technical solutions, advantages, or similar language in the application does not imply that all features and advantages can be achieved in any single embodiment. On the contrary, it can be understood that the description of a feature or advantage means that the specific technical feature, technical solution, or advantage is included in at least one embodiment. Therefore, the description of technical features, technical solutions, or advantages in the specification does not necessarily refer to the same embodiment. Furthermore, the technical features, technical solutions, and advantages described in the embodiments can be combined in any appropriate manner. Those skilled in the art will understand that the embodiments can be implemented without one or more specific technical features, technical solutions, or advantages of a particular embodiment. In other embodiments, additional technical features and advantages can be identified in specific embodiments that do not embody all embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0025] Figure 1 A system architecture diagram of a behavior early warning and digital witnessing system for an evaluation site provided by an embodiment of the present application;
[0026] Figure 2 A flowchart of a behavior early warning and digital witnessing method for an evaluation site provided by an embodiment of the present application;
[0027] Figure 3 A flowchart of another behavior early warning and digital witnessing method for an evaluation site provided by an embodiment of the present application;
[0028] Figure 4 A flowchart of another behavior early warning and digital witnessing method for an evaluation site provided by an embodiment of the present application;
[0029] Figure 5 A flowchart of another behavior early warning and digital witnessing method for an evaluation site provided by an embodiment of the present application;
[0030] Figure 6 A structural diagram of an electronic device provided by an embodiment of the present application;
[0031] Figure 7 A hardware structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0032] In the description of the present application, unless otherwise specified, " / " means "or", for example, A / B can mean A or B. "And / or" in this document is only a description of the relationship between the associated objects, which means that there can be three relationships, for example, A and / or B, which can mean: A exists alone, A and B exist together, and B exists alone. In addition, "at least one" means one or more, and "multiple" means two or more. "First", "second", and the like do not limit the quantity and execution order, and "first", "second", and the like do not necessarily mean different.
[0033] It should be noted that in the present application, "exemplary" or "for example" means to serve as an example, illustration or description. Any embodiment or design scheme described as "exemplary" or "for example" in the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the use of "exemplary" or "for example" is intended to present the relevant concept in a specific manner.
[0034] The behavior early warning and digital witnessing method for an evaluation site provided by an embodiment of the present application can be applied to the behavior early warning and digital witnessing system 100 for an evaluation site as shown in Figure 1 Figure 1 As shown, the communication system comprises a data acquisition device 101 and an electronic device 102.
[0035] The data acquisition device 101 is configured to acquire multi-modal data with time stamps in the bid evaluation site, and the multi-modal data comprises video data, audio data, operation log data, and environmental parameters. The electronic device 102 is configured to extract a feature vector of each type of multi-modal data respectively, and perform data fusion and alignment based on the time stamps to obtain a multi-modal risk feature vector.
[0036] As an example, the data acquisition device 101 includes but is not limited to a monitoring device, a microphone, and a sensor. The monitoring device is configured to acquire the video data. The microphone is configured to acquire the audio data. The sensor is configured to acquire the operation log data.
[0037] To solve the technical problem of insufficient intelligence in the prior art, the embodiment of the present application provides a behavior early warning and digital witnessing method for a bid evaluation site. Figure 2 As shown, the method comprises the following steps.
[0038] S201, acquiring multi-modal data with time stamps in the bid evaluation site.
[0039] The multi-modal data comprises video data, audio data, operation log data, and environmental parameters.
[0040] In a possible implementation, a monitoring device is deployed in the bid evaluation room to acquire video data, a microphone is used to acquire audio data, an interface with a bid evaluation system is used to acquire various operation behavior records of bid evaluators in the bid evaluation system, including login, click, page switching, score submission, etc., and an environmental sensor is used to acquire environmental information such as temperature and humidity in the bid evaluation room, and all data are bound with time stamps.
[0041] As an example, in the embodiment of the present application, a high-definition camera array (8 IP cameras with 1080P resolution and a frame rate of 30 fps) is deployed in the bid evaluation room to capture the behavior features of the bid evaluators. A microphone array (4 omnidirectional microphones with a sampling rate of 16 kHz) is used to collect voice signals and realize sound source positioning. The bid evaluation system interface accesses the operation log of the bid evaluation platform through an API to record user behaviors such as clicking and page switching. Temperature and humidity monitoring devices are used to collect environmental information. All acquisition devices are equipped with high-precision time stamp modules (error ≤1 ms) to ensure accurate alignment of multi-source data in the time dimension.
[0042] S202, the electronic device extracts a feature vector of each type of multi-modal data respectively, and performs data fusion and alignment based on a time stamp to obtain a multi-modal risk feature vector.
[0043] In a possible implementation, the electronic device first extracts behavior features of the video data, the audio data, and the operation log data, and the behavior features include: video behavior features, voice behavior features, and operation behavior features. Then, the electronic device extracts static features of the operation log data and the environmental parameters. The behavior features and the static features are time-aligned on a time axis by using a sliding window to obtain the multi-modal risk feature vector.
[0044] S203, the electronic device inputs the multi-modal risk feature vector into a risk assessment model to determine a risk level of the bid site.
[0045] In a possible implementation, the multi-modal risk feature vector with a time sequence structure is input into the risk assessment model, and modeling analysis is performed from two dimensions of time sequence behavior features and static features. The final risk probability score is obtained according to the importance of the behavior time sequence features and the static features. The risk probability score is compared with a set risk level threshold value, and a risk level is output.
[0046] S204, the electronic device performs risk behavior early warning and evidence storage based on the risk level of the bid site.
[0047] In a possible implementation, a warning level is set, and when the risk level of the bid site reaches a corresponding warning level, a corresponding warning method is taken, such as a pop-up warning, on-site personnel verification, and the like. The video, audio, operation log, and environmental data related to the warning are stored as evidence.
[0048] The method for bid site behavior early warning and witnessing provided in the application improves the intelligent level of bid site behavior early warning and witnessing by deploying multiple sensing devices in the bid site, including a high-definition video camera, a microphone array, an environmental sensor, and an interface of a bid evaluation system, and collecting multi-modal data such as video, audio, operation log, and environmental parameters in real time. The unified time stamp mechanism is combined to achieve millisecond-level synchronization, effectively avoiding the problems of asynchronous multi-source data and fragmented perception, ensuring the time sequence accuracy and semantic integrity of risk behavior analysis.
[0049] In a possible implementation, the method for bid site behavior early warning and witnessing is combined with a risk behavior analysis model. Figure 2 For example, Figure 3 The specific implementation steps of the above S202 of extracting a feature vector of each type of multi-modal data respectively and performing data fusion and alignment based on a time stamp to obtain a multi-modal risk feature vector can be implemented by the following S301-S303, which are described in detail as follows:
[0050] S301, the electronic device extracts behavior features of video data, audio data, and operation log data.
[0051] In a possible implementation, the multi-modal risk feature vector includes behavior timing features and static features. The behavior features include video behavior features, speech behavior features, and operation behavior features. A person outline in the video data is extracted by an image recognition algorithm, and the extracted person mask features and person position coordinates are taken as the video behavior features. Acoustic features in the audio data are extracted by STFT and MFCC, and the acoustic features are taken as the speech behavior features. The target log data is output by cleaning the operation log data, and the target log data is taken as the operation behavior features.
[0052] As an example, in the embodiment of the present application, the video data is denoised by using the OpenCV library, background subtraction is performed by using the MOG2 algorithm, the person outline is recognized by using the YOLOv5 model, the key frames are extracted and converted into the H.264 encoding format, and the video features are output, including the human target bounding box coordinates, the moving range and other information. The audio data is converted into a frequency spectrum by using the short-time Fourier transform, the acoustic features are extracted by using the MFCC (Mel frequency cepstrum coefficient), and the dimension is reduced to 13 dimensions, the audio features are output, including the tone, the intonation, the semantic features and other information. The operation log is cleaned to remove invalid operation records, repeated events and the like, and after being sorted according to the time stamp, the structured JSON format is generated and encoded by using the word embedding, the features with timing in the operation log are extracted, including the operation event sequence, the operation time, the operation frequency and other information.
[0053] S302, the electronic device extracts static features of the operation log data and the environment parameters.
[0054] In a possible implementation, the static behavior features in the operation log are extracted, including the login method, the login terminal, whether a certain type of operation is missing and other information. The environment data is normalized to the [0, 1] interval.
[0055] S303, the electronic device time-aligns the behavior features and the static features on the time axis by using the sliding window method, and obtains a multi-modal risk feature vector.
[0056] In a possible implementation, the space-time alignment engine aligns the behavior features and the static features in time axis through a sliding window based on a timestamp matching mechanism, unifies them to the same time axis, and generates a multi-modal risk feature vector. Specifically, the time axis is initialized; in each time window, the timestamp of all behavior features is traversed to find the timestamp closest to the time window; if the matching is successful, the behavior feature corresponding to the matched timestamp and the static feature are taken as a group of multi-modal risk feature vectors; if the matching fails, the current window is skipped and the next window is entered, and the matching process is repeated.
[0057] As an example, in the embodiment of the application, the sliding window method is used for timestamp matching of multi-modal data, the window length is 5 seconds, the step length is 1 second, and the multi-modal risk feature vector under the unified time axis is generated. First, the time axis T=[t1, t2,..., tn] is initialized, where t i is the timestamp of each modal data; then, for each time window w e T, all modal data is traversed to find the timestamp closest to w; if the matching is successful, the three behavior features corresponding to the timestamp in the time window and the static feature are taken as a group of multi-modal risk feature vectors; if the matching fails, the current window is skipped and the next window is entered.
[0058] In the feature modeling layer, the multi-modal data collected is divided into time sequence behavior features and static features according to functional characteristics, and the sliding window mechanism is used for time alignment and feature fusion. This method not only enhances the ability to capture the time and state of behavior occurrence, but also realizes the alignment and fusion of multi-modal data feature extraction, providing more effective data support for subsequent tasks.
[0059] In a possible implementation, the multi-modal risk feature vector is input into the risk assessment model to determine the risk level of the bid evaluation site. Figure 2 As Figure 4 The above S203 inputs the multi-modal risk feature vector into the risk assessment model to determine the risk level of the bid evaluation site. The specific implementation steps can be realized by the following S401-S404, which are described in detail as follows:
[0060] S401, the electronic device sorts the behavior features in the plurality of multi-modal risk feature vectors based on the time sequence between the plurality of multi-modal risk feature vectors, to obtain a behavior time sequence feature.
[0061] In a possible implementation, first, the plurality of multi-modal risk feature vectors collected in the continuous time period are arranged in time sequence according to the timestamp, a behavior feature sequence with time dependence is constructed, and the behavior-related features in each vector, such as video behavior features, speech features, and operation behavior features, are extracted to form a behavior feature sequence with clear time sequence.
[0062] It should be noted that the continuous behavior time sequence input is the basis for subsequent dynamic modeling to identify behavior evolution patterns. Through time sequencing, dynamic characteristics such as the sequence of abnormal behaviors, trend changes, and repeated patterns can be captured, providing a time sequence basis for determining whether there is a potential violation of rules.
[0063] S402, the electronic device inputs the behavior time sequence feature into the long short-term memory network model, and outputs a dynamic risk probability score.
[0064] In a possible implementation, the behavior time sequence feature sequence is input into a trained long short-term memory (LSTM, Long Short-Term Memory) network. The LSTM network models the dynamic changes of behavior in the time dimension, outputs a dynamic risk prediction score in each time window, and performs standardization processing to output a dynamic risk probability score.
[0065] As an example, in the embodiment of the present application, the dynamic risk probability score P norm satisfies the following formula:
[0066]
[0067] wherein P lstm represents the dynamic risk probability score output by the long short-term memory network in the current window, μ L represents the historical average risk probability score output by the long short-term memory network, and σ L represents the historical standard deviation output by the long short-term memory network.
[0068] S403, the electronic device inputs the static feature into the isolation forest model, and outputs a static anomaly probability score.
[0069] In a possible implementation, the static feature is input into an isolation forest (Isolation Forest) model, an isolated path is constructed based on a random tree structure, and standardization processing is performed to output a static anomaly probability score.
[0070] As an example, in the embodiment of the present application, the static anomaly probability score S norm satisfies the following formula:
[0071]
[0072] wherein S IF is the current static anomaly probability score output by the isolation forest.
[0073] It should be noted that the isolation forest is used to predict the static anomaly probability score, and the purpose is to identify more stable abnormal patterns other than behavior, such as deviation of judges' scoring habits from the group, use of suspicious devices, incomplete operation behavior, etc.
[0074] S404. Electronic equipment weights the dynamic risk probability score and the static anomaly probability score according to their importance, determines the risk probability score, and outputs the risk level at the bid evaluation site.
[0075] In one possible implementation, the two probability scores obtained from the LSTM and Isolation Forest models are weighted and fused to integrate the dynamic risk of behavior and the static anomaly risk, thereby achieving a comprehensive risk assessment of the bidding site. Based on the set risk level threshold, the risk score is mapped to the final risk level.
[0076] As an example, the risk probability score R score Satisfy the following formula:
[0077] R score =α·P norm +β·S nor m
[0078] Where α is the temporal behavior weight, β is the static feature weight, and P norm S represents the standardized time-series risk probability. norm This represents the standardized static anomaly probability.
[0079] As an example, in this embodiment of the application, the risk probability score is compared with the set risk probability threshold, and the risk level is output, wherein the risk level is divided into four levels R∈{0,1,2,3}.
[0080] In terms of risk analysis technology, this application introduces a fusion model architecture, utilizes long short-term memory networks to mine temporal risks of behavior, combines the isolated forest algorithm to identify static anomalies, outputs risk scores from both dynamic and static perspectives, and obtains a comprehensive risk score through a weighted mechanism, thereby improving the accurate identification of complex behavioral combinations.
[0081] In one possible implementation, combining Figure 2 ,like Figure 5 As shown, S204, based on the risk level at the bid evaluation site, conducts risk behavior warnings and evidence preservation, which can be specifically implemented through S501-S503, as detailed below:
[0082] S501. Electronic devices determine the corresponding level of early warning response based on the risk level.
[0083] In one possible implementation, the risk level output above is accepted, and based on a preset warning level configuration table, it is determined whether to trigger a response mechanism and select the corresponding processing method.
[0084] As an example, in the embodiments of the present application, the preset risk level R corresponds to the following response mechanism: R=0: no warning; R=1: control console pop-up prompt; R=2: sound and light prompt in the bid evaluation room; R=3: control console pop-up + sound and light prompt + administrator mobile terminal push.
[0085] S502, when the risk level is greater than or equal to the preset level threshold, the electronic device intercepts the audio data, video data, and operation log data in the first time period.
[0086] In a possible implementation, when the current risk level is greater than or equal to the level threshold, an intercepting instruction is immediately sent to the data cache area, a time window centered on the risk time is traced back and extracted, and data interception in the first time period is formed. The intercepted data includes video data, audio data, and operation log data.
[0087] As an example, in the embodiments of the present application, when the risk level R≥2, the system automatically intercepts the related audio and video segments (length ±5 seconds).
[0088] S503, the electronic device based on the audio data, video data, and operation log data in the first time period.
[0089] The evidence package includes a timestamp, a risk level, a confidence, audio data, video data, and operation log data.
[0090] In a possible implementation, the intercepted multi-modal data is uniformly organized as a structured evidence package, including a risk event timestamp, a current risk level, a confidence output by the risk assessment model, and intercepted original audio, video, and operation log information. The evidence package can be encrypted and stored in a trusted evidence storage system, such as a private blockchain or a digital signature storage platform, to ensure its integrity, non-repudiation, and subsequent verifiability.
[0091] In a possible implementation, after the multi-modal risk feature vector is input into the risk assessment model to determine the risk level of the bid evaluation site, the multi-modal risk feature vector used to determine the risk level of the bid evaluation site is used as training data to train and optimize iterations of the risk assessment model, under the condition that the artificial confirmation of the risk level of the bid evaluation site is consistent with the actual risk level of the bid evaluation site.
[0092] The application combines multi-modal risk perception and dynamic response mechanism, automatically determines whether to trigger an early warning response based on the risk level output by the risk assessment module, implements a hierarchical early warning strategy according to the level difference, realizes instant reminding and manual checking intervention of abnormal behavior, and effectively suppresses the occurrence of irregular behavior. When the risk level reaches the preset threshold, the system automatically traces back and intercepts the multi-modal data before and after the risk moment, forms a data segment in the first time period, covers the core behavior evidence such as audio, video and operation log, and ensures that the key behavior can be traced. And the intercepted data is organized into a structured evidence package and stored in a trusted storage platform with non-tamperable ability, realizing the integrity protection of electronic evidence and improving the intelligentization and credibility level of the behavior supervision in the bid evaluation site.
[0093] The above describes the scheme of the embodiments of the application mainly from the perspective of device implementation. It can be understood that each device, for example, an electronic device, includes at least one of a corresponding hardware structure and a software module for executing each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of the examples described in the embodiments disclosed herein, the application can be realized in the form of hardware or a combination of hardware and computer software. Whether a certain function is driven by hardware or computer software to drive hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the application.
[0094] The embodiments of the application can divide the functional units of the electronic device according to the above method examples. For example, each functional unit can be divided according to each function, or two or more functions can be integrated in one processing unit. The integrated unit can be realized in the form of hardware or software functional unit. It should be noted that the division of units in the embodiments of the application is illustrative, and is only a logical functional division. Actual implementation can have another division method.
[0095] In the case of integrated units, Figure 6 A possible structural schematic diagram of the electronic device (denoted as electronic device 60) involved in the above embodiments is shown, which includes a processing unit 601 and a communication unit 602, and can also include a storage unit 603. Figure 6 The structural schematic diagram shown can be used to illustrate the structure of the electronic device involved in the above embodiments.
[0096] When Figure 6The structure diagram shown is used to show the structure of the electronic device involved in the above embodiment. The processing unit 601 is used to control and manage the action of the electronic device. The communication unit 602 is used for communication between the electronic device and other devices. The storage unit 603 is used to store the program code and data of the electronic device.
[0097] For example, the communication unit 602 is used to obtain time-stamped multi-modal data in the bid evaluation site, and the multi-modal data includes video data, audio data, operation log data, and environmental parameters.
[0098] The processing unit 602 is used to extract the feature vector of each type of multi-modal data respectively, and perform data fusion and alignment based on the time stamp to obtain a multi-modal risk feature vector. The multi-modal risk feature vector is input into a risk assessment model to determine the risk level of the bid evaluation site. Based on the risk level of the bid evaluation site, risk behavior warning and evidence storage are performed.
[0099] In a possible implementation, the processing unit 602 is further configured to: the multi-modal risk feature vector includes behavior time sequence features and static features; the feature vector of each type of multi-modal data is extracted respectively, and data fusion and alignment are performed based on the time stamp to obtain a multi-modal risk feature vector, including: extracting behavior features of video data, audio data, and operation log data, the behavior features including: video behavior features, voice behavior features, and operation behavior features; extracting static features of operation log data and environmental parameters; and the behavior features and the static features are time-aligned through a sliding window on a time axis to obtain the multi-modal risk feature vector.
[0100] In a possible implementation, the processing unit 602 is further configured to extract behavior features based on video data, audio data, and operation log data, including: extracting a person outline in the video data through an image recognition algorithm, and taking the extracted person mask feature and person position coordinates as video behavior features; extracting acoustic features in the audio data through STFT and MFCC, and taking the acoustic features as voice behavior features; cleaning the operation log data to output target log data, and taking the target log data as operation behavior features.
[0101] In a possible implementation, the processing unit 602 is further configured to time-align the behavior features and the static features through a sliding window on a time axis, including: initializing the time axis; in each time window, traversing all time stamps of the behavior features to find a time stamp closest to the time window; if the matching is successful, the behavior feature corresponding to the matched time stamp and the static feature are taken as a group of multi-modal risk feature vectors; and if the matching fails, the current window is skipped, and the next window is entered to repeat the traversal and matching process.
[0102] In a possible implementation, the processing unit 602 is further configured to input the multi-modal risk feature vector into a risk assessment model, determine a risk level of the bid evaluation site, including: based on the time sequence between the multiple sets of multi-modal risk feature vectors, sorting the behavior features in the multiple sets of multi-modal risk feature vectors to obtain a behavior time sequence feature; inputting the behavior time sequence feature into a long short-term memory network model to output a dynamic risk probability score; inputting the static feature into an isolation forest model to output a static anomaly probability score; weighting the dynamic risk probability score and the static anomaly probability score according to importance to determine a risk probability score, and outputting the risk level of the bid evaluation site.
[0103] In a possible implementation, the dynamic risk probability score P norm satisfies the following formula:
[0104]
[0105] wherein P lstm represents the dynamic risk probability score of the current window output by the long short-term memory network, μ L represents the historical average risk probability score output by the long short-term memory network, and σ L represents the historical standard deviation output by the long short-term memory network.
[0106] In a possible implementation, the risk probability score R score satisfies the following formula:
[0107] R score = α·P norm + β·S norm
[0108] wherein α is a time sequence behavior weight, β is a static feature weight, P norm is the standardized time sequence risk probability, and S norm is the standardized static anomaly probability
[0109] In a possible implementation, the processing unit 602 is further configured to, based on the risk level of the bid evaluation site, perform risk behavior early warning and evidence storage, including: based on the risk level, determining a corresponding level of early warning response; when the risk level is greater than or equal to a preset level threshold, intercepting audio data, video data, and operation log data in a first time period; based on the audio data, the video data, and the operation log data in the first time period, storing an evidence package, the evidence package including: a timestamp, a risk level, a confidence, audio data, video data, and operation log data.
[0110] In a possible implementation, the processing unit 602 is further configured to, after inputting the multi-modal risk feature vector into the risk assessment model and determining the risk level of the bid evaluation site, further include: in a case where the artificial confirmation of the risk level of the bid evaluation site is consistent with the actual risk level of the bid evaluation site, inputting the multi-modal risk feature vector used to determine the risk level of the bid evaluation site as training data to iteratively train and optimize the risk assessment model.
[0111] The processing unit 601 can be a processor or a controller, and the communication unit 602 can be a communication interface, a transceiver, a transceiver, a transceiver circuit, a transceiver device, etc. The communication interface is a general term, which can include one or more interfaces. The storage unit 603 can be a memory. When the electronic device 60 is a chip, the processing unit 601 can be a processor or a controller, and the communication unit 602 can be an input interface and / or an output interface, a pin or a circuit, etc. The storage unit 603 can be a storage unit (e.g., a register, a cache, etc.) within the chip, or can be a storage unit (e.g., a read-only memory (ROM), a random access memory (RAM), etc.) located outside the chip.
[0112] The communication unit can also be referred to as a transceiving unit. The antenna and control circuit with transceiving function in the electronic device 60 can be regarded as the communication unit 602 of the electronic device 60, and the processor with processing function can be regarded as the processing unit 601 of the electronic device 60. Optionally, the device for realizing the receiving function in the communication unit 602 can be regarded as a communication unit, and the communication unit is configured to perform the receiving steps in the embodiments of the present application, and the communication unit can be a receiver, a receiver, a receiving circuit, etc. The device for realizing the sending function in the communication unit 602 can be regarded as a sending unit, and the sending unit is configured to perform the sending steps in the embodiments of the present application, and the sending unit can be a transmitter, a sender, a sending circuit, etc.
[0113] Figure 6If the integrated units in the process are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, in essence, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of software products. These computer software products are stored in a storage medium and include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. Storage media for storing computer software products include various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.
[0114] Figure 6 The units in the process can also be called modules; for example, a processing unit can be called a processing module.
[0115] This application also provides a hardware structure diagram of an electronic device (denoted as electronic device 70), see [link to diagram]. Figure 7 The electronic device 70 includes a processor 701, and optionally, a memory 702 connected to the processor 701.
[0116] In the first possible implementation, see Figure 7 The electronic device 70 also includes a transceiver 703. The processor 701, memory 702, and transceiver 703 are connected via a bus. The transceiver 703 is used to communicate with other devices or communication networks. Optionally, the transceiver 703 may include a transmitter and a receiver. The device in the transceiver 703 that implements the receiving function can be considered as a receiver, which is used to perform the receiving steps in the embodiments of this application. The device in the transceiver 703 that implements the transmitting function can be considered as a transmitter, which is used to perform the transmitting steps in the embodiments of this application.
[0117] Based on the first possible implementation method Figure 7 The structural diagram shown can be used to illustrate the structure of the electronic device involved in the above embodiments.
[0118] in, Figure 7 This can also be illustrated by a system chip in an electronic device. In this case, the actions performed by the aforementioned electronic device can be implemented by this system chip; the specific actions performed can be found above and will not be repeated here.
[0119] In the implementation process, each step in the method provided by the embodiment can be completed by the integrated logic circuit of hardware in the processor or the instruction in the form of software. The steps of the method disclosed by the embodiment of the present application can be directly embodied as hardware processor execution completion, or execution completion by hardware and software module combination in the processor.
[0120] The processor in the present application can include but is not limited to at least one of the following: a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a microcontroller unit (MCU), or various types of computing devices running software, such as artificial intelligence processors, each of which can include one or more cores for executing software instructions to perform operations or processing. The processor can be a separate semiconductor chip, or can be integrated with other circuits as a semiconductor chip, for example, it can form a SoC (system on chip) with other circuits such as coding and decoding circuits, hardware acceleration circuits or various bus and interface circuits, or it can be integrated as a built-in processor in the ASIC. The ASIC integrated with the processor can be packaged separately or packaged together with other circuits. In addition to including cores for executing software instructions to perform operations or processing, the processor can further include necessary hardware accelerators, such as field programmable gate arrays (FPGAs), PLDs (programmable logic devices), or logic circuits that implement special logic operations.
[0121] The memory in the embodiment of the present application can include at least one of the following types: read-only memory (ROM) or other types of static storage devices that can store static information and instructions, random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, and electrically erasable programmable read-only memory (EEPROM). In some scenarios, the memory can also be a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage medium or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited to this.
[0122] The embodiment of the present application further provides a computer readable storage medium, comprising instructions which, when executed on a computer, cause the computer to perform any of the above methods.
[0123] The embodiment of the present application further provides a computer program product comprising instructions which, when executed on a computer, cause the computer to perform any of the above methods.
[0124] The embodiment of the present application further provides a chip, comprising a processor and an interface circuit, wherein the interface circuit is coupled with the processor, the processor is configured to execute a computer program or instructions to implement the above method, and the interface circuit is configured to communicate with other modules outside the chip.
[0125] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer readable storage medium or transmitted from one computer readable storage medium to another computer readable storage medium, for example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer readable storage medium can be any available medium that can be accessed by a computer or include one or more data storage devices such as servers, data centers, etc. that can be integrated with the medium. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0126] Although the present application is described herein in conjunction with various embodiments, other variations of the disclosed embodiments can be understood and implemented by those skilled in the art through viewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "one" does not exclude a plurality. A single processor or other unit can implement several functions listed in the claims. Some measures described in mutually different dependent claims can be combined and produce a good result.
[0127] Although the application has been described in conjunction with specific embodiments thereof, it is evident that many alternatives, modifications and variations will be apparent to those skilled in the art. Accordingly, it is intended to embrace all such alternatives, modifications and variations that fall within the scope of the application. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation, as it should be understood that various modifications and equivalents can be used without departing from the spirit and scope of the application. Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense.
Claims
1. A bid evaluation site behavior early warning and digital witnessing method, characterized in that, The method comprises the following steps: acquiring multi-modal data with time stamps in the bid evaluation site, the multi-modal data comprising video data, audio data, operation log data, and environmental parameters; extracting feature vectors of each type of multi-modal data respectively, and performing data fusion and alignment based on the time stamps to obtain multi-modal risk feature vectors; inputting the multi-modal risk feature vectors into a risk assessment model to determine the risk level of the bid evaluation site; based on the risk level of the bid evaluation site, performing risk behavior early warning and evidence storage.
2. The method of claim 1, wherein, The multi-modal risk feature vectors comprise behavior time sequence features and static features; The step of extracting feature vectors of each type of multi-modal data respectively, and performing data fusion and alignment based on the time stamps to obtain multi-modal risk feature vectors comprises: extracting behavior features of the video data, the audio data, and the operation log data, the behavior features comprising video behavior features, voice behavior features, and operation behavior features; extracting static features of the operation log data and the environmental parameters; aligning the behavior features and the static features on a time axis by means of a sliding window to obtain multi-modal risk feature vectors.
3. The method of claim 2, wherein, The step of extracting behavior features based on the video data, the audio data, and the operation log data comprises: extracting a person outline in the video data by means of an image recognition algorithm, and taking the extracted person mask features and person position coordinates as the video behavior features; extracting acoustic features in the audio data by means of STFT and MFCC, and taking the acoustic features as the voice behavior features; cleaning the operation log data to output target log data, and taking the target log data as the operation behavior features.
4. The method of claim 2, wherein, The step of aligning the behavior features and the static features on a time axis by means of a sliding window comprises: initializing a time axis; in each time window, traversing the time stamps of all behavior features to find the time stamp closest to the time window; if the matching is successful, taking the behavior features corresponding to the matched time stamp and the static features as a group of multi-modal risk feature vectors; if the matching fails, skipping the current window, entering the next window, and repeating the traversal and matching process.
5. The method of claim 1, wherein, The step of inputting the multi-modal risk feature vectors into a risk assessment model to determine the risk level of the bid evaluation site comprises: sorting the behavior features in multiple groups of multi-modal risk feature vectors based on the time sequence between the multiple groups of multi-modal risk feature vectors to obtain behavior time sequence features; inputting the behavior time sequence features into a long short-term memory network model to output a dynamic risk probability score; inputting the static features into an isolation forest model to output a static abnormality probability score; weighting the dynamic risk probability score and the static abnormality probability score according to importance to determine a risk probability score, and outputting the risk level of the bid evaluation site.
6. The method of claim 5, wherein, The dynamic risk probability score P norm satisfies the following equation: where P lstm denotes the dynamic risk probability score output by the long short-term memory network for the current window, μ L denotes the historical average risk probability score output by the long short-term memory network, σ L denotes the historical standard deviation output by the long short-term memory network.
7. The method of claim 5, wherein, said risk probability score R score satisfies the following equation: R score = a · P norm + β · S norm wherein a is a time series behavior weight, β is a static feature weight, P norm is a standardized time series risk probability, S norm is a standardized static anomaly probability.
8. The method of claim 1, wherein, The step of performing risk behavior early warning and evidence storage based on the risk level of the bid evaluation site comprises: determining a corresponding level of early warning response based on the risk level. When the risk level is greater than or equal to a preset level threshold, the audio data, the video data, and the operation log data in the first time period are intercepted; Based on the audio data, the video data, and the operation log data in the first time period, an evidence package is stored, and the evidence package includes a timestamp, a risk level, a confidence level, the audio data, the video data, and the operation log data.
9. The method of claim 1, wherein, After the multi-modal risk feature vector is input into a risk assessment model to determine the risk level of the bid site, the method further includes: In a case where the risk level of the bid site is confirmed by a human to be consistent with an actual risk level of the bid site, the multi-modal risk feature vector used to determine the risk level of the bid site is used as training data to train and optimize iterations of the risk assessment model.
10. A behavior early warning and digital witnessing device for bid site, characterized in that, The device includes a communication unit and a processing unit. The communication unit is configured to acquire multi-modal data with a timestamp of a bid site, and the multi-modal data includes video data, audio data, operation log data, and environmental parameters. The processing unit is configured to extract a feature vector of each type of multi-modal data respectively, and perform data fusion and alignment based on the timestamp to obtain a multi-modal risk feature vector. The processing unit is further configured to input the multi-modal risk feature vector into a risk assessment model to determine a risk level of the bid site, and perform a risk behavior warning based on the risk level of the bid site.
Citation Information
Cited By
Stepped dynamic early warning method based on mixed data source
CN121580182A