A subway scene recognition and control method and device based on open source honkong
By using a subway scene recognition method based on the open-source HarmonyOS and leveraging distributed soft bus protocol and multi-level dynamic matching technology, the problem of low accuracy in complex scene recognition in subway entrance and exit control systems has been solved. This enables coordinated control of multiple devices and accurate scene recognition, thereby improving the security and management efficiency of subway entrances and exits.
Patent Information
- Application Number
- CN202610734130.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-26
- Publication Date
- 2026-08-25
AI Technical Summary
The existing subway entrance and exit control system cannot accurately identify complex scenarios, has a high false alarm rate, cannot adapt to the differentiated control needs of different time periods, and has a rigid algorithm model that cannot achieve precise linkage and handling of abnormal events.
A subway scene recognition method based on open-source HarmonyOS is adopted. Multimodal data is acquired through a distributed soft bus protocol, preprocessed, and then subjected to multi-level dynamic matching. Combined with multi-dimensional factors, scene recognition and control are performed to achieve linkage control of multiple devices.
It improved the accuracy of complex scene recognition, reduced the false alarm rate, and achieved accurate recognition and adaptation to different scenarios, ensuring the safety and efficient management of subway entrances and exits.
Smart Images

Figure CN122634005A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology and can be applied to the field of transportation technology, particularly to a method and device for subway scene recognition and control based on the open-source HarmonyOS. Background Technology
[0002] The current development of the subway industry necessitates a focus on identifying abnormal scenarios at subway entrances and exits, which are enclosed spaces, to ensure the safety of equipment and personnel.
[0003] Existing subway entrance and exit control systems rely on general operating systems for scene recognition, supporting only basic card swiping or facial recognition for access. They cannot identify complex scenarios such as tailgating, card breaking, crowd gathering, equipment malfunction, and fire anomalies when entrance and exit roller shutters are opening and closing. For complex scenarios, alarms are triggered based on single events without multi-dimensional factor analysis, resulting in a high false alarm rate and an inability to accurately respond to abnormal events. Furthermore, the algorithm model is rigid, using fixed operating logic, which cannot adapt to the differentiated control needs of different time periods such as morning peak, evening peak, off-peak, and late night, leading to a high false recognition rate and long response delay.
[0004] Therefore, there is an urgent need for a subway scene recognition and control method based on the open-source HarmonyOS that can set dynamic model algorithm matching to adapt to different scene types and improve the accurate recognition of subway entrance and exit scenes. Summary of the Invention
[0005] This invention provides a subway scene recognition and control method and device based on open-source HarmonyOS, which solves the problem of low accuracy in existing complex scene recognition. By setting up an open-source HarmonyOS distributed interconnection architecture, multi-level dynamic matching of scene recognition models is performed to realize distributed linkage control of multiple devices and model algorithm adaptation for different scenarios.
[0006] According to one aspect of the present invention, a subway scene recognition and control method based on the open-source HarmonyOS is provided, comprising: The multimodal data collected by multimodal sensors in the target subway station within a set time period is obtained based on the open-source HarmonyOS distributed soft bus protocol. The multimodal data is preprocessed to obtain a unified feature vector; The unified feature vector and the scene model library are matched in multiple layers to determine the target scene model; Based on the target scene model, scene recognition is performed on the unified feature vector to obtain the target scene; the target scene includes target scene information and target scene level. The target scenario strategy corresponding to the target scenario is determined based on the scenario strategy list and the target scenario level, and the target scenario strategy is executed based on the open-source HarmonyOS distributed soft bus.
[0007] According to another aspect of the present invention, a subway scene recognition and control device based on the open-source HarmonyOS is provided, comprising: The acquisition module is used to acquire multimodal data collected by multimodal sensors in a target subway station within a set time period based on the open-source HarmonyOS distributed soft bus protocol. The preprocessing module is used to preprocess the multimodal data to obtain a unified feature vector; The multi-layer matching module is used to perform multi-layer matching on the unified feature vector and the scene model library to determine the target scene model. The scene recognition module is used to perform scene recognition on the unified feature vector based on the target scene model to obtain the target scene; the target scene includes target scene information and target scene level; The execution module is used to determine the target scenario policy corresponding to the target scenario based on the scenario policy list and the target scenario level, and execute the target scenario policy based on the open-source HarmonyOS distributed soft bus.
[0008] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to execute the subway scene recognition and control method based on open-source HarmonyOS as described in any embodiment of the present invention.
[0009] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions, the computer instructions being configured to cause a processor to execute and implement the subway scene recognition and control method based on open-source HarmonyOS as described in any embodiment of the present invention.
[0010] According to another aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements the subway scene recognition and control method based on open-source HarmonyOS according to any embodiment of the present invention.
[0011] The technical solution of this invention solves the problem of low accuracy in the recognition of complex scenes by setting up an open-source HarmonyOS distributed interconnection architecture and performing multi-level dynamic matching of scene recognition models. This enables distributed linkage control of multiple devices and model algorithm adaptation for different scenarios.
[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 This is a flowchart of a subway scene recognition and control method based on the open-source HarmonyOS, provided according to an embodiment of the present invention; Figure 2 This is a flowchart of a subway scene recognition and control method based on the open-source HarmonyOS, provided according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a subway scene recognition and control device based on the open-source HarmonyOS according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of an electronic device that implements the subway scene recognition and control method based on the open-source HarmonyOS according to the embodiments of the present invention. Detailed Implementation
[0015] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0016] It should be noted that the terms "comprising" and "having" and any variations thereof in the specification, claims and accompanying drawings of this invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such process, method, product or device.
[0017] Furthermore, it should be noted that the information collected in the technical solution of this invention is information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of related data all comply with the relevant laws, regulations and standards of relevant countries and regions, necessary confidentiality measures have been taken, and public order and good morals are not violated. Corresponding operation entry points are provided for users to choose to authorize or refuse.
[0018] Figure 1 This invention provides a flowchart of a subway scene recognition and control method based on the open-source HarmonyOS. This invention is applicable to scene recognition at subway entrances and exits, particularly for subway entrances and exits running the open-source HarmonyOS operating system. This method can be executed by a subway scene recognition and control device based on HarmonyOS, which can be implemented in hardware and / or software and can be configured on a server. Figure 1 As shown, the method includes: S110: Based on the open-source HarmonyOS distributed soft bus protocol, acquire multimodal data collected by multimodal sensors in the target subway station within a set time period.
[0019] Among them, the open-source HarmonyOS distributed soft bus protocol is a protocol-free communication executed by the open-source HarmonyOS operating system installed in the target subway station. It is used to collect multimodal data from multimodal sensors installed in the target subway station. The multimodal sensors may include integrated cameras, access control card readers, temperature and humidity sensors, smoke sensors, roller shutter door status sensors, equipment monitoring sensors and other sensing devices, which are integrated into a super device and deployed at the subway entrances and exits. The multimodal data may include video collected by integrated cameras, personnel passage records collected by access control card readers, temperature and humidity data collected by temperature and humidity sensors, smoke data collected by smoke sensors, roller shutter door data collected by roller shutter door status sensors, and equipment operation data collected by equipment monitoring sensors.
[0020] Specifically, based on the open-source HarmonyOS distributed soft bus protocol, it acquires multimodal data collected by super devices integrated into the entrances and exits of the target subway station within a set time period, such as facial data and behavioral data collected by cameras, card swiping status data collected by access control card readers, humidity data collected by temperature and humidity sensors, smoke data collected by smoke sensors, and device data collected by roller shutter door status sensors.
[0021] In an optional embodiment of the present invention, the multimodal data further includes station attribute data uploaded by the subway station and time data corresponding to a set time period. The station attribute data may include station type and station importance level; for example, the station type may include whether the target subway station is a transfer station or a hub station; the station importance level may be the area level near the target subway station, such as being close to government or tourist attractions, which requires special attention to identify the danger level and can reduce the warning weight for large passenger flows to avoid unnecessary waste of resources; the time data is the current time.
[0022] S120. Preprocess the multimodal data to obtain a unified feature vector.
[0023] Preprocessing can involve denoising, normalizing, and feature extraction of the data; the unified feature vector is the feature vector corresponding to each multimodal data with unified dimensions after normalization.
[0024] Specifically, multi-source heterogeneous data is collected uniformly through the open-source HarmonyOS distributed soft bus protocol, and the data is processed by denoising, normalization, and feature extraction to generate feature vectors with unified dimensions.
[0025] Optionally, the multimodal data can be preprocessed to obtain a unified feature vector, including: Feature extraction is performed on facial and behavioral data from multimodal data to obtain facial features and human behavioral features; Feature extraction is performed on the card swipe status data in the multimodal data to obtain the passage record features; Feature extraction is performed on site attribute information and device data in multimodal data to obtain site attribute features and device status features; Feature extraction is performed on time data in multimodal data to obtain time period features; The facial features, personnel behavior features, access record features, site attribute features, device status features, and time period features are normalized to obtain a unified feature vector.
[0026] Among them, facial features are used to identify whitelisted or blacklisted personnel; personnel behavior features are used to identify personnel behavior, such as tailgating or trespassing; access record features are card swiping status, which can be card swiped or not; site attribute features are used to identify site type and site importance level; device status features are whether the gate is abnormal, device CPU utilization, and network latency; time period features indicate time type, such as morning peak, evening peak, or off-peak.
[0027] Specifically, facial features can be extracted from multimodal data using convolutional neural networks; behavioral features can be extracted from multimodal data by combining temporal features; access record features can be extracted from card swipe status data within the current time period using sequence models; site attribute information in multimodal data can be vectorized into text using natural language processing techniques to extract key attributes and obtain site attribute features; device status features can be extracted from device data by combining time series; time data sequence decomposition in multimodal data can obtain time period features; denoising preprocessing can be performed on facial features, behavioral features, access record features, site attribute features, device status features, and time period features; and dimensionality normalization can be performed on the preprocessed facial features, behavioral features, access record features, site attribute features, device status features, and time period features to obtain a unified feature vector, which can be represented as F=[f1,f2,...,fn].
[0028] Understandably, by pre-extracting features from multimodal data and normalizing them, a feature vector with uniform dimensions is obtained, which improves the efficiency of subsequent multi-dimensional dynamic matching of scene models. Furthermore, feature preprocessing reduces the amount of local computation, improves the efficiency of scene recognition, and enables timely control of abnormal scenes, effectively ensuring the safety of personnel in abnormal scenarios at subway entrances and exits.
[0029] S130. Perform multi-level matching on the unified feature vector and scene model library to determine the target scene model.
[0030] The scenario model library stores at least one candidate scenario model, which is matched with a specific scenario type. The candidate scenario model can be a cross-modal model. Each candidate scenario model is stored in correspondence with a specific scenario type. For example, the safety cross-modal model is used for identifying abnormal roller shutter doors, people staying, breaking in, climbing over, or tailgating, and disaster early warning; the environment cross-modal model is used for identifying water level exceeding limits, slippery weather, and abnormal temperature / brightness; the safety cross-modal model is used for identifying large passenger flows, key personnel, equipment failures, and entrance / exit congestion; and the target scenario model is a candidate scenario model used for identifying the current scenario.
[0031] Specifically, based on the unified feature vector and candidate scene models in the scene model library, multi-level matching is performed. For example, hash retrieval can be used for coarse matching, and then fine matching can be performed based on feature vector similarity to determine the target scene model from the candidate scene models.
[0032] In one optional embodiment of the present invention, the candidate scene models in the scene model library have a one-to-one correspondence with the scene types, which is used to identify specific scenes, such as peak period scenes, tailgating scenes, large passenger flow scenes, checkpoint violation scenes, equipment failure scenes, or abnormal behavior scenes, etc.; the candidate scene models can be obtained by training the general model multimodal Transformer model on specific sample scenes, which is used to identify specific scenes in a targeted manner, effectively improving the accuracy and efficiency of scene recognition.
[0033] S140. Based on the target scene model, perform scene recognition on the unified feature vector to obtain the target scene; the target scene includes target scene information and target scene level.
[0034] The target scenario information can include target scenario type, target personnel type, time type, and station type, which are used to determine the scenario level from multiple dimensions. The target scenario type can be a normal scenario, an abnormal scenario, or a high-risk scenario; the target personnel type can be blacklisted personnel, whitelisted personnel, or irrelevant personnel; the time type can be peak operation, off-peak operation, or end of operation; the station type can be a hub station / transfer station, a key station, or an ordinary station; the target scenario level can be divided into five levels, such as level 5 corresponding to extremely high risk, level 4 corresponding to high risk, level 3 corresponding to relatively high risk, level 2 corresponding to general risk, or level 1 corresponding to low risk.
[0035] Specifically, based on the target scene model, the unified feature vector is used to identify the scene and obtain the target scene information corresponding to the scene, such as the target scene type, target personnel type, time type, and site type; and further, the scene level corresponding to the current target scene is determined based on the target scene type, target personnel type, time type, and site type.
[0036] Optionally, scene recognition is performed on the unified feature vector based on the target scene model to obtain the target scene, including: Scene identification is performed on the unified feature vector based on the target scene model to obtain target scene information; the target scene information includes target scene type, target personnel type, time type, and site type; The target scene information is mapped based on the scene information mapping table to obtain multi-dimensional scene factor attribute values; The scenario risk value is obtained by weighting and aggregating the attribute values of multi-dimensional scenario factors. The target scenario level is determined based on the scenario risk value and risk threshold.
[0037] The scenario information mapping table represents the risk value attribute factors corresponding to each dimension of the target scenario information. Multi-dimensional scenario factor attribute values can include scenario factors, personnel factors, time period factors, and site attribute factors. Scenario risk values are numerical values obtained using the fuzzy comprehensive evaluation method based on multi-dimensional scenario factors, ranging from 0 to 100, and are used to represent the risk level of the current scenario. Risk thresholds can be threshold ranges corresponding to risk levels, such as 0-20 for level 1 risk, 20-40 for level 2 risk, 40-60 for level 3 risk, 60-80 for level 4 risk, and 80-100 for level 5 risk. Specifically, based on the target scenario model, scene identification is performed on the unified feature vector to obtain the target scenario type, target personnel type, time type, and site type; the target scenario information is mapped according to the scenario information mapping table to obtain multi-dimensional scenario factor attribute values; the scenario factor in the multi-dimensional scenario factor attribute values is determined according to the target scenario type and the scenario type mapping table, such as a normal scenario with a scenario factor of 10 points; an abnormal scenario with a scenario factor of 30-70 points; and a high-risk scenario with a scenario factor of 80-100 points; the personnel factor in the multi-dimensional scenario factor attribute values is determined according to the target personnel type and the personnel type mapping table, such as a target personnel type being suspected of terrorism or blacklisted personnel with a personnel factor of 40 points; and a target personnel type being whitelisted or internal personnel with a personnel factor of 0 points; If the personnel type is unknown, the personnel factor is 15 points. The time period factor in the multi-dimensional scenario factor attribute values is determined according to the time type and time period type mapping table. For example, if the time type is peak operation, the time period factor is 10 points; if the time type is off-peak, the time period factor is 5 points; if the time type is end of operation, the time period factor is 20 points. The station attribute factor in the multi-dimensional scenario factor attribute values is determined according to the station type and station type mapping table. For example, if the station type is a hub station or transfer station, the station attribute factor is 30 points; if the station type is a key station, the station attribute factor is 20 points; if the station type is a regular station, the station attribute factor is 10 points. The scenario factor, personnel factor, time period factor, and station attribute factor in the multi-dimensional scenario factor attribute values are weighted and aggregated to obtain the scenario risk value. The calculation method for the scenario risk value is as follows: in, This represents the risk value for the scenario. A can be set to [0.5, 0.25, 0.15, 0.1]. The value can be U=[U1,U2,U3,U4]; U1 is the scene factor; U2 is the personnel factor; U3 is the time period factor; and U4 is the site factor.
[0038] The target scenario level is determined based on the scenario risk value and risk threshold. For example, a scenario risk value of 0-20 indicates low risk and its target scenario level is 1; a scenario risk value of 20-40 indicates moderate risk and its target scenario level is 2; a scenario risk value of 40-60 indicates relatively high risk and its target scenario level is 3; a scenario risk value of 60-80 indicates high risk and its target scenario level is 4; and a scenario risk value of 80-100 indicates extremely high risk and its target scenario level is 5.
[0039] Understandably, by identifying target scenarios through scenario models and making multi-dimensional risk assessments based on scenario type, personnel type, time type, and site type, and setting up a multi-factor linkage risk classification mechanism, the accuracy of high-risk scenario identification has been improved by 59%, the false alarm rate has been reduced by 78%, and the workload of security personnel in handling ineffective situations has been significantly reduced. This enables targeted identification of scenarios based on site type, and effectively achieves precise linkage processing in the future.
[0040] S150. Determine the target scenario policy corresponding to the target scenario based on the scenario policy list and the target scenario level, and execute the target scenario policy based on the open-source HarmonyOS distributed soft bus.
[0041] The scenario policy list stores the scenario policies corresponding to each scenario level; it can record target scenario information or generate alarm information based on target scenario information for early warning processing.
[0042] Specifically, the target scenario strategy is determined based on the scenario strategy list and the target scenario level. For example, if the target scenario level is Level 1 risk, the strategy is to log the target scenario information and upload it to the cloud platform. If the target scenario level is Level 2 risk, the strategy is to generate alarm information based on the target scenario information and push the alarm information to the HarmonyOS smart terminal of the site security personnel. If the target scenario level is Level 3 risk, the strategy is to lock the corresponding gate, trigger on-site audio and visual alarms, generate alarm information based on the target scenario information, and push the alarm information to the site control room. If the target scenario level is Level 4 risk, the strategy is to determine the type of target personnel in the target scenario information, link surrounding cameras to automatically track the target, broadcast a prompt within the site, and notify the site security personnel to handle the situation on-site. If the target scenario level is Level 5 risk, the strategy is to directly trigger a 110 alarm, push the on-site personnel images collected from the multimodal data corresponding to the set time period to the public security department, broadcast an evacuation prompt throughout the site, and close all entrance and exit gates. This allows the open-source HarmonyOS distributed soft bus linkage device to execute the target scenario strategy. This invention, through its embodiment, acquires multimodal data collected by multimodal sensors in a target subway system within a set time period using the open-source HarmonyOS distributed soft bus protocol. The multimodal data is preprocessed to obtain a unified feature vector. Multi-level matching is performed between the unified feature vector and a scene model library to determine the target scene model. Scene recognition is then performed on the unified feature vector based on the target scene model to obtain the target scene. The target scene includes target scene information and a target scene level. The corresponding target scene strategy is determined based on the scene strategy list and the target scene level, and the target scene strategy is executed based on the open-source HarmonyOS distributed soft bus. This technical solution, by setting up an open-source HarmonyOS distributed interconnection architecture and performing multi-level dynamic matching of scene recognition models, achieves distributed linkage control of multiple devices and adapts model algorithms to different scenes, thus solving the problem of low accuracy in existing complex scene recognition.
[0043] Figure 2 This is a flowchart illustrating a subway scene recognition and control method based on the open-source HarmonyOS, according to an embodiment of the present invention. This embodiment supplements the multi-layer matching method of the target scene model based on the above embodiments. It should be noted that for parts not detailed in this embodiment, please refer to the relevant descriptions in other embodiments. Figure 2 As shown, the method includes: S210: Based on the open-source HarmonyOS distributed soft bus protocol, acquire multimodal data collected by multimodal sensors in the target subway station within a set time period.
[0044] S220. Preprocess the multimodal data to obtain a unified feature vector.
[0045] S230. Perform multi-dimensional operations on the unified feature vector based on the weight factor adjustment rule to obtain the target feature vector.
[0046] Among them, the weight factor adjustment rule is a rule for allocating weight factors to multi-dimensional features in a unified feature vector. It is used to dynamically adjust the time period weight factor, type weight factor, and health weight factor. The multi-dimensional calculation can be calculated using the Hadamard product. The target feature vector is the normalized feature weight vector.
[0047] Specifically, the weight factors of the features corresponding to each multimodal data in the unified feature vector are allocated according to the multidimensional feature weight factor allocation rule in the weight factor adjustment rule, and the target feature vector is obtained by multidimensional operation based on the Hadamard product to calculate the weight factors and the unified feature vector.
[0048] Optionally, a multi-dimensional operation is performed on the unified feature vector based on the weight factor adjustment rule to obtain the target feature vector, including: Based on the traffic type mapping table in the weight factor adjustment rule, the time period features and personnel behavior features in the unified feature vector are weighted and adjusted to obtain the time period weight factor. Based on the site type mapping table in the weight factor adjustment rule, the weights of the site attribute features and access record features in the unified feature vector are adjusted to obtain the type weight factor. Based on the device mapping table in the weight factor adjustment rule, the weights of device status features and facial features in the unified feature vector are adjusted to obtain the health weight factor. The target feature vector is obtained by calculating the unified feature vector based on the time period weight factor, type weight factor, and health degree weight factor.
[0049] The access type mapping table stores the weight allocation rules for time period features and personnel behavior features; the station type mapping table stores the weight allocation rules for station attribute features and access record features; and the device mapping table stores the weight allocation rules for device status features and facial features.
[0050] Specifically, based on the weighting rules for time period characteristics and personnel behavior characteristics in the traffic type mapping table within the weighting factor adjustment rules, the weights of personnel behavior characteristics are adjusted to obtain the time period weighting factor: for example, during morning and evening peak hours, the weight of abnormal behavior characteristics is reduced by 20%; during off-peak hours, the weight of abnormal behavior characteristics is increased by 25%, etc.; based on the weighting rules for station attribute characteristics and traffic record characteristics in the station type mapping table within the weighting factor adjustment rules, the weights of traffic record characteristics are adjusted to obtain the type weighting factor: for example, if the station attribute characteristic is a transfer station or hub station, the weight of the traffic record characteristic for high passenger flow is increased by 30%; if the station attribute characteristic is a regular station... If the passage record shows tailgating or checkpoint trespassing, the weight of these features increases by 20%. If the station attribute feature is a key station, the weight of the passage record shows dangerous goods identification or anti-terrorism features increasing by 40%. Based on the weight allocation rules for device status features and facial features in the device mapping table in the weight factor adjustment rules, the weight of facial features is adjusted to obtain the health weight factor. If the device status feature is a CPU utilization rate greater than 85% or a network latency greater than 250ms, the weight of facial features and personnel behavior features is reduced. The Hadamard product of the time period weight factor, type weight factor, and health weight factor on the unified feature vector is calculated to obtain the target feature vector.
[0051] Understandably, by setting up a three-dimensional dynamic weight allocation, the weights can be dynamically adjusted. By using three dimensions—time period weight, site type, and device health—the system avoids triggering scenario judgments with a single data point, which could lead to misjudgments. This approach can adapt to the differentiated management needs of different time periods, such as morning peak, evening peak, off-peak, and late at night, ensuring the targeted selection of the scenario recognition model.
[0052] S240. Based on the target feature vector, perform hash retrieval on the candidate scene models in the scene model library to determine at least two scene models to be matched.
[0053] Among them, hash retrieval can be a 64-bit locality-sensitive hash fast retrieval method; the scenario model to be matched can be the candidate scenario model selected in the coarse matching stage.
[0054] Specifically, a hash operation is performed on the target feature vector, and the Hamming distance is calculated between it and the candidate scene models in the scene model library. At least two scene models to be matched are determined from the candidate scene models based on the hash distance threshold.
[0055] Optionally, a hash search is performed on candidate scene models in the scene model library based on the target feature vector to determine at least two scene models to be matched, including: The Hamming distance between the target feature vector and the model features corresponding to the candidate scene models in the scene model feature library is calculated to obtain the distance hash value; At least two scene models to be matched are determined from the candidate scene models based on the distance threshold and the distance hash value.
[0056] Among them, the distance hash value is the value obtained by calculating the Hamming distance between each candidate scene model and the target feature vector; the distance threshold is used to filter relevant models, preferably 5.
[0057] Specifically, the target feature vector is subjected to local sensitive hashing to obtain the target hash value. The target hash value is then compared with the model features corresponding to the candidate scene models in the scene model feature library to calculate the Hamming distance, thus obtaining the distance hash value between the target feature vector and each candidate scene model. Based on the distance threshold and the distance hash value, the candidate scene models corresponding to the distance hash value less than the distance threshold are selected as the scene models to be matched.
[0058] Understandably, using a 64-bit locality-sensitive hashing fast retrieval method to perform coarse matching on candidate scene models can filter out more than 90% of irrelevant models, with a matching time of <10ms, thus improving the efficiency of scene model selection.
[0059] S250. Based on the target feature vector and the model features of the scene model to be matched, perform similarity matching to determine the target scene model.
[0060] Specifically, the similarity between the target feature vector and the model features of the scene model to be matched is calculated, and the scene model to be matched with the higher similarity value is taken as the target scene model.
[0061] Optionally, similarity matching is performed based on the target feature vector and the model features of the scene model to be matched to determine the target scene model, including: The similarity between the target feature vector and the model features of the scene model to be matched is calculated to obtain the cosine similarity value; The target feature vector and the model features of the scene model to be matched are calculated using a weighted Euclidean distance to obtain the distance value; A fused similarity value is obtained by performing a fusion calculation based on the distance scaling factor, cosine similarity value, and distance value. The target scene model is determined from the scene models to be matched based on the fusion of similarity values and similarity thresholds.
[0062] Among them, the cosine similarity value is obtained by calculating the cosine similarity between the target feature vector and the model features of the scene model to be matched; the distance value is obtained by calculating the weighted Euclidean distance; the distance scaling factor is preferably 10; and the similarity threshold is a fused similarity threshold used to determine the target scene model.
[0063] Specifically, the similarity between the target feature vector and the model features of the scene model to be matched is calculated to obtain a cosine similarity value; a weighted Euclidean distance is calculated between the target feature vector and the model features of the scene model to be matched to obtain a distance value; and a fused similarity value is obtained by fusing the distance scaling factor, the cosine similarity value, and the distance value. The formula for calculating the fused similarity value is as follows: in, To fuse similarity values; For dynamic coefficients; Cosine similarity value; This is the distance value; This is the distance scaling factor.
[0064] Based on the fused similarity value and the similarity threshold, the target scene model is determined from the scene models to be matched. If there is a fused similarity value greater than the similarity threshold, the scene model to which the fused similarity value belongs is taken as the target scene model. If there are at least two fused similarity values greater than the similarity threshold, manual verification is performed. If all fused similarity values are less than the similarity threshold, the target feature vector is adjusted based on the preset adjustment rules. The above steps are repeated until there is at least one fused similarity value greater than the similarity threshold.
[0065] Understandably, by setting up fusion similarity calculation to perform fine matching on the scene model to be matched, and finding the target scene model suitable for the target feature vector for scene recognition based on the similarity threshold, the hierarchical matching mechanism ensures that the recognition accuracy is improved by 47% and the matching speed is improved by 75%.
[0066] S260. Based on the target scene model, perform scene recognition on the unified feature vector to obtain the target scene; the target scene includes target scene information and target scene level.
[0067] S270. Determine the target scenario policy corresponding to the target scenario based on the scenario policy list and the target scenario level, and execute the target scenario policy based on the open-source HarmonyOS distributed soft bus.
[0068] This invention performs model matching based on a unified feature vector and candidate scene models in a scene model library. It adopts a multi-dimensional dynamic adaptation hierarchical model matching algorithm to generate the optimal algorithm model that is adapted to the current subway scene, effectively realizing full-scene perception of subway entrances and exits and dynamic algorithm adaptation, thus solving the problem of poor algorithm model adaptability in existing subway entrance and exit scene recognition.
[0069] Figure 3 This invention provides a schematic diagram of a subway scene recognition and control device based on the open-source HarmonyOS. This invention is applicable to scene recognition at subway entrances and exits, particularly for subway entrance and exit scene recognition using the open-source HarmonyOS operating system. This subway scene recognition and control device based on HarmonyOS can be implemented in hardware and / or software, and can be configured on a server. Figure 3 As shown, the subway scene recognition and control device 300 based on the open-source HarmonyOS includes an acquisition module 310, a preprocessing module 320, a multi-layer matching module 330, a scene recognition module 340, and an execution module 350. The acquisition module 310 is used to acquire multimodal data collected by multimodal sensors in the target subway station within a set time period based on the open-source HarmonyOS distributed soft bus protocol. The preprocessing module 320 is used to preprocess multimodal data to obtain a unified feature vector; The multi-layer matching module 330 is used to perform multi-layer matching on the unified feature vector and the scene model library to determine the target scene model. The scene recognition module 340 is used to perform scene recognition on a unified feature vector based on a target scene model to obtain the target scene; the target scene includes target scene information and target scene level. The execution module 350 is used to determine the target scenario policy corresponding to the target scenario based on the scenario policy list and the target scenario level, and execute the target scenario policy based on the open source HarmonyOS distributed soft bus.
[0070] This invention, through its embodiment, acquires multimodal data collected by multimodal sensors in a target subway station within a set time period using the open-source HarmonyOS distributed soft bus protocol. The multimodal data is preprocessed to obtain a unified feature vector. Multi-level matching is performed between the unified feature vector and a scene model library to determine the target scene model. Scene recognition is then performed on the unified feature vector based on the target scene model to obtain the target scene. The target scene includes target scene information and a target scene level. The target scene strategy corresponding to the target scene is determined based on the scene strategy list and the target scene level, and the target scene strategy is executed based on the open-source HarmonyOS distributed soft bus. This technical solution, by setting up an open-source HarmonyOS distributed interconnection architecture and performing multi-level dynamic matching of scene recognition models, achieves distributed linkage control of multiple devices and adapts model algorithms to different scenes, thus solving the problem of low accuracy in existing complex scene recognition.
[0071] Optionally, the multi-layer matching module 330 also includes a multi-dimensional operation unit, a hash retrieval unit, and a similarity matching unit; A multi-dimensional operation unit is used to perform multi-dimensional operations on a unified feature vector based on a weight factor adjustment rule to obtain a target feature vector. The hash retrieval unit is used to perform hash retrieval on candidate scene models in the scene model library based on the target feature vector, and determine at least two scene models to be matched. The similarity matching unit is used to perform similarity matching between the target feature vector and the model features of the scene model to be matched, and to determine the target scene model.
[0072] Optionally, the multi-dimensional operation unit is also used to adjust the weights of time period features and personnel behavior features in the unified feature vector based on the pass type mapping table in the weight factor adjustment rule, so as to obtain the time period weight factor. Based on the site type mapping table in the weight factor adjustment rule, the weights of the site attribute features and access record features in the unified feature vector are adjusted to obtain the type weight factor. Based on the device mapping table in the weight factor adjustment rule, the weights of device status features and facial features in the unified feature vector are adjusted to obtain the health weight factor. The target feature vector is obtained by calculating the time period weight factor, type weight factor, and health degree weight factor on the unified feature vector.
[0073] Optionally, the hash retrieval unit is also used to calculate the Hamming distance between the target feature vector and the model features corresponding to the candidate scene models in the scene model feature library to obtain the distance hash value; At least two scene models to be matched are determined from the candidate scene models based on the distance threshold and the distance hash value.
[0074] Optionally, the similarity matching unit is also used to calculate the similarity between the target feature vector and the model features of the scene model to be matched, and obtain the cosine similarity value. The target feature vector and the model features of the scene model to be matched are calculated using a weighted Euclidean distance to obtain the distance value; A fused similarity value is obtained by performing a fusion calculation based on the distance scaling factor, cosine similarity value, and distance value. The target scene model is determined from the scene models to be matched based on the fusion of similarity values and similarity thresholds.
[0075] Optionally, the preprocessing module 320 is also used to extract features from face data and behavior data in multimodal data to obtain face features and human behavior features; Feature extraction is performed on the card swipe status data in the multimodal data to obtain the passage record features; Feature extraction is performed on site attribute information and device data in multimodal data to obtain site attribute features and device status features; Feature extraction is performed on time data in multimodal data to obtain time period features; The facial features, personnel behavior features, access record features, site attribute features, device status features, and time period features are normalized to obtain a unified feature vector.
[0076] The optional scene recognition module 340 is also used to perform scene recognition on the unified feature vector based on the target scene model to obtain target scene information; the target scene information includes target scene type, target personnel type, time type and site type; The target scene information is mapped based on the scene information mapping table to obtain multi-dimensional scene factor attribute values; The scenario risk value is obtained by weighting and aggregating the attribute values of multi-dimensional scenario factors. The target scenario level is determined based on the scenario risk value and risk threshold.
[0077] The subway scene recognition and control device based on open-source HarmonyOS provided in this embodiment of the invention can execute the subway scene recognition and control method based on open-source HarmonyOS provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0078] According to embodiments of the present invention, the present invention also provides an electronic device, a readable storage medium, and a computer program product.
[0079] Figure 4A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0080] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0081] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0082] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, central processing unit (CPU), graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as the subway scene recognition and control method based on the open-source HarmonyOS.
[0083] In some embodiments, the subway scene recognition and control method based on the open-source HarmonyOS can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the subway scene recognition and control method based on the open-source HarmonyOS described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to execute the subway scene recognition and control method based on the open-source HarmonyOS by any other suitable means (e.g., by means of firmware).
[0084] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include: implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0085] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0086] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0087] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0088] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0089] A computing system can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product within the cloud computing service system. This addresses the shortcomings of traditional physical hosts and dedicated virtual services, such as high management difficulty and weak business scalability.
[0090] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0091] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A subway scene recognition and control method based on open-source HarmonyOS, characterized in that, include: The multimodal data collected by multimodal sensors in the target subway station within a set time period is obtained based on the open-source HarmonyOS distributed soft bus protocol. The multimodal data is preprocessed to obtain a unified feature vector; The unified feature vector and the scene model library are matched in multiple layers to determine the target scene model; Based on the target scene model, scene recognition is performed on the unified feature vector to obtain the target scene; the target scene includes target scene information and target scene level. The target scenario strategy corresponding to the target scenario is determined based on the scenario strategy list and the target scenario level, and the target scenario strategy is executed based on the open-source HarmonyOS distributed soft bus.
2. The method according to claim 1, characterized in that, The step of performing multi-level matching on the unified feature vector and the scene model library to determine the target scene model includes: Based on the weight factor adjustment rule, the unified feature vector is subjected to multi-dimensional operations to obtain the target feature vector; Based on the target feature vector, perform hash retrieval on the candidate scene models in the scene model library to determine at least two scene models to be matched. The target scene model is determined by performing similarity matching based on the target feature vector and the model features of the scene model to be matched.
3. The method according to claim 2, characterized in that, The step of performing multi-dimensional operations on the unified feature vector based on the weight factor adjustment rule to obtain the target feature vector includes: Based on the traffic type mapping table in the weight factor adjustment rule, the time period features and personnel behavior features in the unified feature vector are weighted and adjusted to obtain the time period weight factor. Based on the site type mapping table in the weight factor adjustment rule, the weights of the site attribute features and access record features in the unified feature vector are adjusted to obtain the type weight factor. Based on the device mapping table in the weight factor adjustment rule, the weights of device status features and facial features in the unified feature vector are adjusted to obtain the health weight factor. The target feature vector is obtained by calculating the unified feature vector based on the time period weight factor, the type weight factor, and the health weight factor.
4. The method according to claim 2, characterized in that, The step of performing hash retrieval on candidate scene models in the scene model library based on the target feature vector to determine at least two scene models to be matched includes: The Hamming distance between the target feature vector and the model features corresponding to the candidate scene models in the scene model feature library is calculated to obtain the distance hash value. At least two scene models to be matched are determined from the candidate scene models based on the distance threshold and the distance hash value.
5. The method according to claim 2, characterized in that, The step of performing similarity matching based on the target feature vector and the model features of the scene model to be matched, to determine the target scene model, includes: The similarity between the target feature vector and the model features of the scene model to be matched is calculated to obtain the cosine similarity value; A weighted Euclidean distance is calculated between the target feature vector and the model features of the scene model to be matched to obtain a distance value; A fused similarity value is obtained by performing a fusion calculation based on the distance scaling factor, the cosine similarity value, and the distance value. Based on the fused similarity value and similarity threshold, the target scene model is determined from the scene models to be matched.
6. The method according to claim 1, characterized in that, The preprocessing of the multimodal data to obtain a unified feature vector includes: Feature extraction is performed on the facial data and behavioral data in the multimodal data to obtain facial features and human behavioral features; Feature extraction is performed on the card swipe status data in the multimodal data to obtain passage record features; Feature extraction is performed on the site attribute information and device data in the multimodal data to obtain site attribute features and device status features; Feature extraction is performed on the time data in the multimodal data to obtain time period features; The facial features, personnel behavior features, access record features, site attribute features, device status features, and time period features are subjected to dimensionality normalization processing to obtain a unified feature vector.
7. The method according to claim 1, characterized in that, The step of performing scene recognition on the unified feature vector based on the target scene model to obtain the target scene includes: Based on the target scene model, scene recognition is performed on the unified feature vector to obtain target scene information; the target scene information includes target scene type, target personnel type, time type, and site type; The target scene information is mapped based on the scene information mapping table to obtain multi-dimensional scene factor attribute values; The multi-dimensional scenario factor attribute values are weighted and aggregated to obtain the scenario risk value; The target scenario level corresponding to the target scenario type is determined based on the scenario risk value and risk threshold.
8. A subway scene recognition and control device based on open-source HarmonyOS, characterized in that, include: The acquisition module is used to acquire multimodal data collected by multimodal sensors in a target subway station within a set time period based on the open-source HarmonyOS distributed soft bus protocol. The preprocessing module is used to preprocess the multimodal data to obtain a unified feature vector; The multi-layer matching module is used to perform multi-layer matching on the unified feature vector and the scene model library to determine the target scene model. The scene recognition module is used to perform scene recognition on the unified feature vector based on the target scene model to obtain the target scene; the target scene includes target scene information and target scene level; The execution module is used to determine the target scenario policy corresponding to the target scenario based on the scenario policy list and the target scenario level, and execute the target scenario policy based on the open-source HarmonyOS distributed soft bus.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to execute the subway scene recognition and control method based on open source HarmonyOS as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the subway scene recognition and control method based on open-source HarmonyOS as described in any one of claims 1-7.
11. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the subway scene recognition and control method based on open-source HarmonyOS according to any one of claims 1-7.