Multimedia monitoring and management system for Internet of Things of venue
Through the IoT multimedia monitoring and management system, the audience data is collected and analyzed in real time, and the explanation model is dynamically adjusted, which solves the problem that personalized explanation cannot be provided in the existing technology, and improves the user experience and product promotion effect.
Patent Information
- Application Number
- CN202510829036.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-20
AI Technical Summary
The existing venue guide technology cannot provide personalized explanation services based on the real-time characteristics of the audience, resulting in poor user experience and affecting product sales.
Through the Internet of Things multimedia monitoring and management system, the audience's voice interaction and mobile trajectory data are collected in real time, the basic population categories are analyzed and matched, and the speech speed, volume, language style and content depth of the explanation model are dynamically adjusted to meet the needs of different groups of people.
The precise docking of commentary services has been realized, which has improved users' acceptance of commentary content, improved user experience and product promotion efficiency, and reduced resource waste.
Smart Images

Figure CN120337043A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multimedia monitoring, and specifically relates to an Internet of Things multimedia monitoring and management system for venues. Background Art
[0002] Smart venues are the product of the deep integration of traditional venues and digital technologies, and their core lies in intelligent perception, data-driven decision-making, and user-friendly services. Through technological iteration, smart venues are evolving from single-functional places into comprehensive service platforms that integrate safety, efficiency, and experience. In the future, with the popularization of AIoT technology, smart venues will play a more important role in smart cities.
[0003] In the field of venue operation and management, the Internet of Things multimedia monitoring and management system, relying on advanced technologies such as sensors and network communication, efficiently completes data collection and device linkage, and has become a key tool for the operation and management of many venues, providing comprehensive data insights for venue operation.
[0004] Existing venue guidance technologies, such as intelligent robots, mostly operate based on preset programs or rely on users' active selection, lacking real-time pertinence and being difficult to meet the personalized commentary needs of different audiences. In contrast, the data advantages of the Internet of Things multimedia monitoring and management system can effectively make up for this shortcoming. The system collects real-time audience characteristic data through sensors. After analysis and processing, the information is transmitted to the corresponding multimedia language device, which then calls the commentary model that suits the current user, and then realizes the switching and output of the corresponding commentary content; In this context, how to achieve the precise docking of the commentary service with the real-time characteristics of the audience based on the commentary model switching mechanism of the Internet of Things multimedia monitoring and management system, and further prompt the corresponding multimedia language device to provide more targeted commentary services for stayers with different characteristics has become an urgent problem to be solved. Summary of the Invention
[0005] The purpose of the present invention is to provide an Internet of Things multimedia monitoring and management system for venues, which solves the technical problem of dynamically switching commentary models for different groups of people to improve acceptance.
[0006] An Internet of Things multimedia monitoring and management system for venues includes: A behavior perception module that collects behavioral characteristic data of the target population, where the behavioral characteristic data at least includes voice interaction data and movement trajectory data, and the language interaction data includes preparatory words and execution words; A data analysis module that analyzes and processes the collected behavioral characteristic data to extract key characteristic information that can reflect the acceptance degree of the target population for the commentary model; The model call module matches the target population with the preset population classifications according to the key feature information, determines the basic population category to which the target population belongs, and calls the corresponding basic explanation model for the basic population category; The data comparison module continuously monitors the real-time behavior feature data of the target population during the explanation process and compares it with the expected behavior feature data under the called basic explanation model; The model adjustment module, when the difference between the real-time behavior feature data and the expected behavior feature data exceeds the preset threshold, dynamically adjusts the current explanation model from one or more dimensions of the explanation speech rate, volume, language style, and content depth according to the difference type to generate an adapted current explanation model to improve the user's acceptance of the explanation content.
[0007] As a further solution of the present invention: the preset population classifications at least include the elderly, middle-aged people, and young people. The basic explanation model corresponding to the elderly is a model with a slow speech rate and a loud voice. The basic explanation model corresponding to the middle-aged people is a model with a moderate speech rate and refined content. The basic explanation model corresponding to the young people is a model with a fast speech rate and a lively language style.
[0008] As a further solution of the present invention: the analysis and processing of the collected behavior feature data include: Performing speech recognition and semantic analysis on the voice interaction data, and extracting the voice response duration, question frequency, semantic understanding accuracy rate, preparation words, execution words, and switching words as key feature information; Analyzing the mobile trajectory data and extracting the stay time, moving speed, and moving direction change frequency as key feature information; Analyzing the operation instruction data and extracting the instruction repetition times, instruction execution success rate, and instruction issuing interval time as key feature information.
[0009] As a further solution of the present invention: the model adjustment module further includes a preliminary switching module and a switching operation module: The preliminary switching module, when obtaining the voice interaction data of the target population, if a preparation word is detected, determines the age classification result of the current target population, pre-loads the explanation model that meets the current age classification in the preset database in advance based on the age classification result, and starts a timer for countdown; The switching operation module, during the timing process, continuously obtains the key feature data of the target population. If a key behavior action or execution word is obtained, immediately calls the explanation model that matches the current age classification for explanation; if no key behavior action or execution word is obtained after the timing ends, uses the preparation word as a loading word to call the explanation model related to the loading word in the preset database for loading.
[0010] As a further solution of the present invention: during the explanation process, continuously monitor the real-time behavior characteristic data of the target population, and compare it with the expected behavior characteristic data under the called basic explanation model. Specifically: Establish a corresponding relationship database between the basic explanation model and the expected behavior characteristic data; Real-time collect the behavior characteristic data of the target population, and match it with the expected behavior characteristic data corresponding to the basic explanation model called in the corresponding relationship database; Calculate the difference degree between the real-time behavior characteristic data and the expected behavior characteristic data, and judge whether the difference degree exceeds the preset threshold.
[0011] As a further solution of the present invention: when multiple preparatory words are obtained, perform the following steps: Determine whether the various preparatory words point to the same meaning, and calculate whether the proportion of the preparatory words pointing to the same meaning exceeds the preset value; If the proportion of the preparatory words pointing to the same meaning is greater than the preset value, continue to execute the preloading by calling the corresponding explanation model in the preset database in advance; If the proportion of the preparatory words pointing to the same meaning is less than or equal to the preset value, continuously obtain whether other preparatory words continue to appear within the preset time, and determine in real time within the preset time whether it exceeds the preset value after the newly added preparatory words are added. If it exceeds, continue to execute the preloading by calling the corresponding explanation model in the preset database in advance. If it does not exceed, use the filtered preparatory words as the loading words for preloading.
[0012] As a further solution of the present invention: using the filtered preparatory words as the loading words for preloading specifically means: Use the basic requirement words among the multiple preparatory words as the main loading words, and other words as parameter words; Select the explanation model corresponding to the main loading word as the basic explanation method, and adjust the basic explanation method in combination with the corresponding explanation model parameters in the parameter words.
[0013] As a further solution of the present invention: it further includes a derivative recommendation module, and the specific working process is: When the data analysis module obtains preparatory words from the voice interaction data, automatically generate a set of potential demand words related to the preparatory words; Store the set of potential demand words as a preloading item in the temporary buffer area, and start the demand monitoring timer; During the timing of the demand monitoring timer, continuously monitor the voice interaction data of the target population; If an execution word that matches any word in the set of potential demand words appears in the voice interaction data of the target population, the derivative product explanation model corresponding to the execution word is retrieved through the model call module, and the derivative product explanation model is fused with the current basic explanation model and recommended to the target population; If no matching execution word is detected when the demand monitoring timer times out, the preloaded items in the temporary buffer are cleared.
[0014] As a further solution of the present invention: during the timing of the demand monitoring timer, if a specific behavior combination of the target population is detected, derivative recommendation is accelerated: when the user stays in front of a certain product for more than a preset time and a preparation word appears in the voice interaction, the derivative product explanation model is immediately retrieved without waiting for the demand monitoring timer to end.
[0015] As a further solution of the present invention: it further includes a storage module, which is used to store preset population classifications, basic explanation models, a corresponding relationship database between the basic explanation models and expected behavior feature data, and the corresponding relationship between the preparation words and the explanation models.
[0016] Compared with the prior art, the beneficial effects of the present invention are: The present invention adjusts the current explanation model from one or more dimensions of the explanation speech rate, volume, language style, and content depth according to the difference type by collecting the voice interaction and mobile trajectory data of customers in real time, so as to accurately analyze the needs of different populations and generate a personalized explanation model that meets the current population. Then, it can continuously monitor the customer's reaction during the explanation process, which not only improves the customer's shopping experience and makes it easier for them to understand product information, but also helps to promote products efficiently, reduce ineffective resource investment, and optimize the explanation model through the accumulated data, making the explanation service more and more accurate and suitable for different types of populations. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a schematic diagram of the system framework structure of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] The technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0019] In various venues, the venue space is usually divided according to the actual situation, and different commodities and exhibits are placed in different areas. Since different people have different interests in the exhibits in each area of the venue, based on industry common sense, in terms of tour guide and interpretation services, smart venues often use technical means such as intelligent robots, AR / VR tour guide systems, and voice interaction devices to reduce the dependence on manual interpretation. The common functions of multimedia monitoring in venues are mainly real-time monitoring and recording relevant evidence, etc. However, at present, most multimedia monitoring systems are independent systems and are not deeply integrated with the tour guide service. They lack the function of analyzing the characteristics of the people staying in the venue and then enabling the corresponding multimedia language devices to provide more targeted explanations according to the characteristics of different stayers.
[0020] Specifically, existing venue tour guide technologies (such as intelligent robots, AR / VR devices) mostly rely on preset programs or user initiative selection and cannot provide personalized interpretation services according to the real-time characteristics of stayers (such as age, interest, staying duration). For example, when different users show interest in the current commodities and exhibits, existing multimedia devices generally use a single language style for explanation. For the elderly with poor hearing, if the speech rate is too fast and the voice is too low, they may not be able to hear clearly and thus choose to leave; for young people, if the speech rate is too slow and the voice is too loud, they may find it troublesome and leave, which will affect the overall experience effect of customers.
[0021] Obviously, if personalized interpretation services cannot be provided according to the characteristics of different people, the staying time of users will be significantly shortened, which will undoubtedly have an adverse impact on the sales of commodities.
[0022] Please refer to Figure 1 , this application provides a venue Internet of Things multimedia monitoring and management system, including: A behavior perception module that collects the behavior characteristic data of the target population, and the behavior characteristic data at least includes voice interaction data and movement trajectory data, where the language interaction data includes a preparation word and an execution word; A data analysis module that analyzes and processes the collected behavior characteristic data and extracts key characteristic information that can reflect the acceptance degree of the target population for the interpretation model; A model call module that matches the target population with the preset population classification according to the key characteristic information, determines the basic population category to which the target population belongs, and calls the corresponding basic interpretation model for the basic population category; A data comparison module that continuously monitors the real-time behavior characteristic data of the target population during the interpretation process and compares it with the expected behavior characteristic data under the called basic interpretation model; The model adjustment module dynamically adjusts the current commentary model from one or more dimensions including commentary speed, volume, language style, and content depth according to the type of difference when the difference between the real-time behavior feature data and the expected behavior feature data exceeds a preset threshold, and generates an adapted current commentary model to improve user acceptance of the commentary content.
[0023] As an optional embodiment, during the explanation process, the real-time behavior characteristic data of the target group is continuously monitored and compared with the expected behavior characteristic data under the called basic explanation model, specifically: Establish a database of correspondence between basic interpretation models and expected behavioral characteristic data; Collect the behavioral characteristic data of the target population in real time and match it with the expected behavioral characteristic data corresponding to the basic interpretation model called in the corresponding relationship database; Calculate the difference between the real-time behavior feature data and the expected behavior feature data, and determine whether the difference exceeds a preset threshold.
[0024] As an optional embodiment, the preset population classification includes at least the elderly, middle-aged people, and young people. The basic interpretation model corresponding to the elderly is a model with slow speaking speed and loud voice, the basic interpretation model corresponding to the middle-aged people is a model with moderate speaking speed and refined content, and the basic interpretation model corresponding to the young people is a model with fast speaking speed and lively language style.
[0025] The target group’s behavioral characteristics data is used to track the target group’s action path (movement trajectory) in the exhibition area with the help of cameras, Wi-Fi positioning, Bluetooth beacons and other technologies. For example, through the camera’s video analysis technology, the user’s stay time in front of different product exhibits and other data are recorded to determine the user’s interest points; Among them, a microphone array or intelligent voice equipment is used to capture the voice information of the target group in the product display area in real time, and the collected data is pre-processed by noise reduction, format conversion, etc. The above methods are all existing technologies, and other methods can also be used to achieve them, which will not be elaborated here.
[0026] Among them, preparatory words refer to the preparatory language used by users before expressing their needs, such as "I want to know" and "this product", which are used to identify the user's intention to interact; execution words are keywords that clearly express needs or feedback, such as "function", "price", "user experience", etc.
[0027] Among them, natural language processing (NLP) technology is used to perform semantic analysis and sentiment analysis on voice interaction data, and the core needs of users’ questions are understood through semantic analysis.
[0028] Among them, by combining the display area layout, the movement trajectory data is analyzed. If the user stays in front of a certain commodity for more than a preset time value, it indicates that the user has strong interest in the commodity. The above analysis results are converted into quantitative indicators, such as the frequency of voice interaction, the duration of staying in the movement trajectory, etc., as the key feature information reflecting the user's acceptance degree.
[0029] Among them, according to the common user types in the commodity display scenario, multiple basic population categories are preset. According to the behavioral characteristic data information, the current stayers are divided into three types: the elderly, the middle-aged, and the young. And these divisions are not absolute, and should be flexibly adjusted according to the specific situation in actual applications.
[0030] Among them, machine learning classification algorithms (such as support vector machines, random forests) are used, and the key feature information is used as the input to train the model to achieve automatic matching of population categories.
[0031] Among them, after determining the basic population category to which the target population belongs, the model call module retrieves the corresponding basic explanation model from the preset database stored in the cloud server for preloading to reduce the user's waiting time. The preset database can be managed using a relational database (such as MySQL), and fast retrieval and call are achieved through efficient SQL query statements. And multiple preset models can be set in advance according to the actual situation and stored in the preset database, which will not be elaborated here too much.
[0032] Among them, a corresponding relationship database is constructed through historical data accumulation and experimental testing.
[0033] Among them, by collecting the behavioral characteristic information of the target population in real time and matching it with the expected behavioral characteristic data in the corresponding relationship database, each basic explanation model is pre-set with corresponding expected behavioral characteristics that match it; Furthermore, for numerical data (such as voice response duration, movement speed), the absolute difference method is used to calculate the difference value. Specifically: ; for numerical data (such as voice response duration, movement speed), the absolute difference method is used to calculate the difference value: , the comprehensiveness is calculated using the weighted average method to calculate the overall difference degree between the real-time behavioral characteristic data and the expected behavioral characteristic data; and according to the importance of each characteristic index, different weights are assigned. The formula is: , n represents the number of behavioral characteristic indicators participating in the calculation of the overall difference degree, and i is an index variable from 1 to n, which is used to sequentially refer to each specific behavioral characteristic indicator, represents the proportion of the importance of the i-th behavioral characteristic indicator in the calculation of the overall difference degree, refers to the difference measurement value between the real-time data of the i-th behavioral characteristic indicator and the expected data under the corresponding basic explanation model; Furthermore, when the calculated overall difference exceeds a preset threshold, the data comparison module sends a difference warning signal to the model adjustment module to trigger the model adjustment process; if it does not exceed the threshold, the current explanation model is continued to be maintained.
[0034] After receiving the difference warning signal sent by the data comparison module, the model adjustment module first deeply analyzes the difference data to determine the difference type.
[0035] The core of the present invention is: according to the difference type, the current explanation model is adjusted from one or more dimensions of the explanation speech rate, volume, language style, and content depth. If it is determined to be a problem of the explanation speech rate, it is adjusted by modifying the speech rate parameter of the speech synthesis engine. The explanation volume is changed by adjusting the volume output parameter of the audio playback device (such as setting the volume percentage in the multimedia playback software) or the volume parameter of the speech synthesis engine; the corresponding template is retrieved from the preset language style template library to replace the current language style. The language style template library includes various style templates such as professional, popular, and lively. According to the user's comprehension ability and interest level, different levels of content are extracted from the commodity knowledge graph and recombined. For users with relatively weak comprehension ability, the introduction of complex principles is reduced and actual application cases are increased.
[0036] In this way, it is possible to collect the voice interaction and movement trajectory data of customers in real time. For example, when hearing preparation words such as "introduce" or recording the stay time in front of a certain commodity, it is possible to accurately analyze the needs of different groups of people and generate a personalized explanation model that conforms to the current group of people. Then, it is possible to continuously monitor the customer's reaction during the explanation process. For example, if it is found that the customer frequently asks questions, the content depth is adjusted to make the explanation more suitable for their acceptance level. In this way, not only the shopping experience of customers is improved, making it easier for them to understand commodity information, but also it helps to promote commodities efficiently, reduce ineffective resource investment, and optimize the explanation model through the accumulated data, making the explanation service more and more accurate and suitable for different types of people.
[0037] As an optional embodiment, the analysis and processing of the collected behavioral characteristic data include: Performing speech recognition and semantic analysis on the voice interaction data, and extracting the voice response duration, question frequency, semantic understanding accuracy rate, preparation words, execution words, and switching words as key feature information; Analyzing the movement trajectory data, and extracting the stay time, movement speed, and movement direction change frequency as key feature information; Analyzing the operation instruction data, and extracting the instruction repetition times, instruction execution success rate, and instruction issuance interval time as key feature information.
[0038] The present application further proposes to analyze and process the collected behavioral feature data. Specifically, first, a microphone array (such as a 4-microphone circular array) deployed in the commodity display area is used to collect voice signals, and then a speech recognition engine (such as iFlytek offline SDK) is used to convert the analog audio into a text stream; Subsequently, natural language processing (NLP) technology is adopted for semantic analysis to extract voice response duration, question frequency, semantic understanding accuracy rate, preparation words, execution words, and switching words; For example, the time difference from the end of the system playing the explanatory voice to the user's opening response is calculated to obtain the voice response duration, the number of questions per minute is counted as the question frequency, and the semantic understanding accuracy rate is calculated by calculating the cosine similarity between the user's response content and the standard answer in the commodity knowledge base; The extraction of preparation words / execution words is performed by matching a preset keyword library through regular expressions, and the switching words are the words that capture the change of the user's intention, etc. The above are all existing technologies and will not be elaborated too much; To analyze the mobile trajectory data, a UWB positioning system (such as Decawave DW1000 chip) can be used to collect the real-time coordinates of the user, with a positioning accuracy of 10 cm, and the mobile trajectory is recorded at a frequency of 5 Hz, so as to obtain parameters such as residence time, moving speed, and moving direction change frequency. This is an existing technology and will not be elaborated too much; Among them, in a display device that supports touch interaction (such as an intelligent shopping guide screen), the user operation instructions are captured through an event listening mechanism to obtain the instruction repetition times, instruction execution success rate, and instruction issuing interval time. This is an existing technology and will not be elaborated too much.
[0039] As an optional embodiment, the model adjustment module further includes a preliminary switching module and a switching operation module: The preliminary switching module, when obtaining the voice interaction data of the target population, if a preparation word is detected, determines the age classification result of the current target population, preloads the explanatory model that conforms to the current age classification in the preset database in advance based on the age classification result, and starts a timer for countdown; The switching operation module, during the timing process, continuously obtains the key feature data of the target population. If a key behavioral action or execution word is obtained, immediately calls the explanatory model that matches the current age classification for explanation; if no key behavioral action or execution word is obtained after the timing ends, the preparation word is used as the loading word to call the explanatory model related to the loading word in the preset database for loading.
[0040] The present application further proposes that in the commodity display scenario, when the behavior perception module obtains the voice interaction data of the target population, the preliminary switching module first monitors the data in real time. Once a preparation word (such as "I want to know" "Introduce to me") is detected, the age classification process is immediately started; First, the classification process relies on the key feature information extracted by the data analysis module rather than simply dividing by age, such as voice response time, word usage habits, movement trajectory, etc. as division parameters. The specific division method is: for example, if the user's voice response is slow, the wording is simple and direct, and the movement speed is slow, the system will judge it as an elderly person through a machine learning algorithm (such as a support vector machine); if the user's language expression is concise, the question is professional, and the movement speed is moderate, it will be judged as a middle-aged person; if the user's language is lively, frequently uses Internet hot words, and the movement trajectory shows fast browsing characteristics, it will be judged as a young person. Compared with simply dividing by age, this division method is more in line with the selection of appropriate explanation models for explanation, and can better improve users' acceptance of product explanations. For example, a user who is classified as a young person in terms of age, but in reality, has a slow voice response, simple and direct wording, and a slow movement speed, it is incorrect to use the basic explanation model of young people at this time; Secondly, after determining the classification result, the pre-switching module retrieves the explanation model that meets the current classification from the preset database based on the result for pre-loading. The preset database stores basic explanation models designed for different classification groups, such as a model with slow speech speed and loud voice for the elderly, a model with moderate speech speed and concise content for the middle-aged, and a model with fast speech speed and lively language style for the young. The pre-loading operation stores the model data in memory in advance, shortening the response time of subsequent calls; At the same time, the preparatory switching module starts the timer to count down. The countdown duration can be flexibly set according to actual business needs, generally between 5-15 seconds.
[0041] Finally, during the countdown of the timer, the switching operation module will continue to obtain key feature data of the target population. These data come from the behavior perception module and the data analysis module, including the execution words in the voice interaction, the dwell time and direction changes in the movement trajectory, the operation instruction data, etc. If key behavioral actions (such as standing still to observe a product for a long time, repeatedly touching the product) or action words (such as "tell me more" or "demonstrate") are obtained during the timing process, the switching operation module will immediately call the explanation model that matches the current age category for explanation. For example, if it is determined to be a young person and the action word "demonstrate the latest function" is detected, the basic explanation model corresponding to young people will be quickly called to introduce the innovative functions of the product in detail in a lively language style; If no key action or execution word is obtained after the timing ends, the switching operation module will use the preparatory word as the loading word and retrieve and load the corresponding explanation model related to the loading word from the preset database. For example, if the preparatory word is "I want to know", the system will retrieve and load a general product introduction model from the database, which includes basic information, core selling points, etc. of the product, and provide preliminary explanation services for users first, which is beneficial to avoiding user loss during the waiting process, and is also beneficial to ensuring that the system can continuously and accurately provide personalized product explanation services for the target population, improving the user's acceptance of the explanation content and shopping experience.
[0042] As an optional embodiment, when multiple preparatory words are obtained, the following steps are performed: Determine whether the meanings of the preparatory words point to the same meaning, and calculate whether the proportion of the number of preparatory words pointing to the same meaning exceeds the preset value; If the proportion of the number of preparatory words pointing to the same meaning is greater than the preset value, continue to retrieve and preload the corresponding explanation model in the preset database in advance; If the proportion of the number of preparatory words pointing to the same meaning is less than or equal to the preset value, continuously obtain whether other preparatory words continue to appear within the preset time, and determine in real time whether it exceeds the preset value after the newly added preparatory words are added within the preset time. If it exceeds, continue to retrieve and preload the corresponding explanation model in the preset database in advance. If it does not exceed, use the filtered preparatory words as the loading word for preloading.
[0043] This application further proposes that when the behavior perception module obtains the voice interaction data of the target population, if multiple preparatory words are detected (such as "I want to know", "introduce it to me"), first use the word vector model (such as Word2Vec, BERT) in natural language processing (NLP) technology to convert each preparatory word into a vector representation in a high-dimensional space, and calculate their distances in the semantic space through the cosine similarity algorithm to determine whether they point to the same meaning; Determine that the preparatory words with similarity higher than the threshold point to the same meaning. Subsequently, calculate the proportion of the number of preparatory words pointing to the same meaning in the total number of all preparatory words; if the proportion of the number of preparatory words pointing to the same meaning exceeds the preset value, it means that the user's needs are clear and concentrated. At this time, the system immediately executes the operation of the original preparatory switching module: according to the age classification result of the target population (determined by analyzing features such as voice intonation, word usage habits, and movement trajectories), retrieve and preload the corresponding explanation model that meets this age classification in the preset database in advance. For example, if it is determined that the user is middle-aged and the preparatory words all point to "understanding the product cost performance", then preload the refined explanation model corresponding to middle-aged people, highlighting information such as price and performance; When the proportion does not exceed the preset value, the system starts the time window monitoring mechanism and continuously listens for the appearance of new preparatory words within the preset time. During this period, for each newly added preparatory word, the proportion of preparatory words pointing to the same meaning is recalculated in real time, and it is judged whether it exceeds the preset value. For example, if the proportion of the initial 3 preparatory words is only 60%, but 2 more preparatory words pointing to the same meaning are added within 8 seconds, increasing the proportion to 80%, the preloading process is immediately triggered.
[0044] If the proportion still does not exceed the preset value after the preset time ends, it indicates that the user's needs are relatively scattered or ambiguous. At this time, the preparatory words are screened and integrated, and the general interpretation model related to these loading words in the preset database is retrieved for preloading, briefly covering the basic information of all aspects of the commodity to provide a preliminary interpretation for the user; An intelligent screening logic is added to the preloading decision-making link after detecting the preparatory words. It responds quickly when the user's needs are clear and flexibly responds when the needs are ambiguous, which is beneficial to avoiding model mismatches caused by misjudging the needs, thus significantly improving the user experience and the system response efficiency.
[0045] As an optional embodiment, preloading the screened preparatory words as loading words specifically means: Taking the basic requirement words among multiple preparatory words as the main loading words and other words as parameter words; Selecting the interpretation model corresponding to the main loading word as the basic interpretation method and adjusting the basic interpretation method in combination with the corresponding interpretation model parameters in the parameter words.
[0046] Specifically, preloading the screened preparatory words as loading words means taking the basic requirement words among multiple preparatory words as the main loading words and other words as parameter words. By preloading key interpretation resources into memory or related modules in advance, it is possible to respond quickly during user interaction. The main loading words represent the core needs of the user, while the parameter words provide additional personalized information; Secondly, select the interpretation model corresponding to the main loading word as the basic interpretation method. The basic interpretation model is preset by the system to meet the basic needs of most users. For example, in the interpretation system, the basic interpretation model may be a general interpretation template for introducing the basic information of exhibits, such as the name, material, and use of the exhibits; Adjust the basic interpretation method in combination with the corresponding interpretation model parameters in the parameter words. The information in the parameter words (such as language style, speech rate, volume, etc.) is used to perform personalized adjustments to the basic interpretation model. For example, if the parameter word contains "elderly people", the system may adjust the speech rate of the interpretation model to be slower and the volume to be moderate to adapt to the hearing needs of the elderly; if the parameter word contains "young people", the speech rate may be adjusted to be faster and the volume to be normal to meet the preferences of young people.
[0047] As an optional embodiment, it further includes a derivative recommendation module, and the specific workflow is as follows: When the data analysis module obtains a preparatory word from the voice interaction data, it automatically generates a set of potential demand words related to the preparatory word; The set of potential demand words is stored in the temporary buffer as a preloaded item, and a demand monitoring timer is started; During the timing of the demand monitoring timer, continuously monitor the voice interaction data of the target population; If an execution word that matches any word in the set of potential demand words is detected in the voice interaction data of the target population, the derivative product explanation model corresponding to the execution word is retrieved through the model call module, and the derivative product explanation model is fused with the current basic explanation model and recommended to the target population; If no matching execution word is detected after the demand monitoring timer finishes timing, the preloaded items in the temporary buffer are cleared.
[0048] Specifically, when the data analysis module identifies a preparatory word (such as "this mobile phone", "this set of furniture") from the voice interaction data, the semantic expansion algorithm is immediately started. Based on the commodity knowledge graph and natural language processing technology, this algorithm generates a set of potential demand words; First, it extends by combining the core functions of the commodity. Secondly, it recommends associated products according to the commodity usage scenarios. Finally, it automatically supplements competitor comparison words for popular commodities. The generated set of potential demand words is sorted by weight through the TF-IDF algorithm, and the 5-10 most relevant words are selected; After the set of potential demand words is generated, the system stores it in the memory-based temporary buffer (such as Redis cache), and at the same time starts the demand monitoring timer (default setting is 15 seconds). This buffer adopts the LRU (Least Recently Used) elimination strategy to ensure efficient data storage and fast retrieval; During the timing process, the behavior perception module collects voice interaction data at a frequency of 10 times per second, and compares the user's voice with the potential demand words in the buffer through the keyword real-time matching algorithm (such as the AC automaton algorithm). For example, when the user says "Can it filter formaldehyde", the system quickly matches the "purification effect" in the set of potential demand words and triggers the subsequent recommendation logic; Once a matching execution word is detected, the model invocation module immediately retrieves the corresponding explanation model from the preset database of derivative products. This preset database pre-designs differential explanation strategies for different product categories and demand scenarios. The retrieved derivative product explanation model and the current basic explanation model (the main model classified based on user age) are fused through a weight assignment algorithm. For example, if the user is a young person, the basic explanation model accounts for 60% (maintaining a lively language style), and the derivative model accounts for 40% (highlighting professional parameters), finally forming a mixed explanation content of "younger expression + in-depth function analysis". The system converts the fused content into smooth speech through a speech synthesis engine (such as iFlytek's multi-style synthesis) and simultaneously presents graphic and text auxiliary information on the display screen; If no matching execution word is detected when the demand monitoring timer ends, the system automatically clears the temporary buffer area and releases memory resources. When the pre-loaded basic explanation model has not been triggered, the monitoring results of the derivative recommendation module can be used as a supplementary basis for adjusting the pre-loaded model. For example, if "cost performance" frequently appears in the set of potential demand words, the priority of the basic explanation model for middle-aged people can be raised in advance, and the data of unmatched potential demand words will be fed back to the data analysis module for optimizing the subsequent potential demand word generation algorithm; Through the above process, the derivative recommendation module realizes the in-depth mining from the user's explicit needs to implicit needs. Without interrupting the normal explanation process, it improves the accuracy of product recommendation with lightweight monitoring and intelligent fusion strategies, and at the same time avoids user resistance caused by over-recommendation, ultimately enhancing the user's efficiency in obtaining product information and purchase intention.
[0049] As an optional embodiment, during the timing of the demand monitoring timer, if a specific behavior combination of the target population is detected, the derivative recommendation is accelerated: when the user stays in front of a certain product for more than the preset time and a preparation word appears in the voice interaction, the derivative product explanation model is immediately retrieved without waiting for the demand monitoring timer to end.
[0050] Specifically, during the operation of the demand monitoring timer, the system simultaneously monitors the movement trajectory data and voice interaction data of the target population to form a dual triggering condition. The user's location is tracked in real time through the UWB positioning system. When the user's stay duration in a certain product display area exceeds the preset threshold, this behavior is marked as a person with in-depth attention, and the microphone array simultaneously captures voice data. Once the preparation word is recognized, the subsequent process is triggered. Only when both of the above two conditions are met will the system choose to skip the waiting link of the demand monitoring timer and directly start the derivative recommendation; Secondly, by obtaining the triggering of a specific behavior combination, the set of potential demand words related to the preparation words is immediately retrieved (adopting the original generation logic), and the pre-loaded derivative product explanation model is quickly extracted from the temporary cache. For example, when the user stays in the mobile phone exhibition area for 3 minutes and then asks "this mobile phone", the system instantly loads the explanation content corresponding to potential demands such as "performance parameters" and "battery life"; Compared with the conventional timing trigger, the derivative recommendation triggered by the specific behavior combination has a higher priority. The system will temporarily increase the weight of the derivative model in the fusion strategy to ensure that the core information that the user may be interested in is preferentially displayed. For example, the processor performance and battery capacity of the mobile phone are preferentially explained, while the language style characteristics of the basic explanation model are retained. Thus, through the cross-verification of the stay time and voice commands, random browsing behaviors are effectively filtered, which is beneficial to focusing on users with real needs and avoiding ineffective recommendations.
[0051] As an optional embodiment, it further includes a storage module, which is used to store the preset population classification, basic explanation model, the corresponding relationship database between the basic explanation model and the expected behavior feature data, and the corresponding relationship between the preparation words and the explanation model.
[0052] Specifically, the storage module adopts a distributed storage architecture, combines a relational database (such as MySQL) and a non-structured database (such as MongoDB) to implement data classification management. Storing data through the storage module facilitates subsequent calls, which is an existing mature technology and will not be elaborated here.
[0053] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. An Internet of Things multimedia monitoring and management system for a venue, characterized in that, include: A behavior perception module collects behavior characteristic data of the target population, wherein the behavior characteristic data includes at least voice interaction data and movement trajectory data, wherein the language interaction data includes preparation words and execution words; The data analysis module analyzes and processes the collected behavioral feature data to extract key feature information that can reflect the target population's acceptance of the explanation model; The model calling module matches the target population with the preset population classification according to the key feature information, determines the basic population category to which the target population belongs, and calls the corresponding basic interpretation model for the basic population category; The data comparison module continuously monitors the real-time behavioral characteristic data of the target group during the explanation process and compares it with the expected behavioral characteristic data under the called basic explanation model; The model adjustment module dynamically adjusts the current commentary model from one or more dimensions including commentary speed, volume, language style, and content depth according to the type of difference when the difference between the real-time behavior feature data and the expected behavior feature data exceeds a preset threshold, and generates an adapted current commentary model to improve user acceptance of the commentary content.
2. The multimedia monitoring and management system for the venue Internet of Things according to claim 1, characterized in that The preset population classification includes at least the elderly, middle-aged people, and young people. The basic interpretation model corresponding to the elderly is a model with slow speech speed and loud voice, the basic interpretation model corresponding to the middle-aged people is a model with moderate speech speed and refined content, and the basic interpretation model corresponding to the young people is a model with fast speech speed and lively language style.
3. The multimedia monitoring and management system for the venue Internet of Things according to claim 1, characterized in that, The analyzing and processing of the collected behavior characteristic data includes: Perform speech recognition and semantic analysis on the speech interaction data, and extract speech response duration, question frequency, semantic understanding accuracy, preparation words, execution words, and switching words as key feature information; Analyze the movement trajectory data and extract the dwell time, movement speed, and movement direction change frequency as key feature information; The operation instruction data is analyzed to extract the number of instruction repetitions, instruction execution success rate, and instruction issuance interval as key feature information.
4. The multimedia monitoring and management system for the Internet of Things in a venue according to claim 3, wherein, The model adjustment module also includes a preparatory switching module and a switching operation module: The preparation switching module, when acquiring the voice interaction data of the target population, determines the age classification result of the current target population after detecting the preparation words, retrieves the explanation model that meets the current age classification in the preset database in advance based on the age classification result, and starts the timer to count down; Switch the operation module. During the timing process, continuously obtain the key feature data of the target population. If key behavioral actions or execution words are obtained, immediately call the explanation model that matches the current age classification for explanation; if key behavioral actions or execution words are not obtained after the timing ends, the preparation words are used as loading words to load the explanation model related to the loading words in the preset database.
5. The multimedia monitoring and management system for venue Internet of Things according to claim 1, wherein, During the explanation process, the real-time behavior characteristic data of the target group is continuously monitored and compared with the expected behavior characteristic data under the called basic explanation model, specifically: Establish a database of correspondence between basic interpretation models and expected behavioral characteristic data; Collect the behavioral feature data of the target population in real time and match it with the expected behavioral feature data corresponding to the basic interpretation model called from the corresponding relationship database. Calculate the difference degree between the real-time behavioral feature data and the expected behavioral feature data, and judge whether the difference degree exceeds the preset threshold.
6. The multimedia monitoring and management system for the Internet of Things in a venue according to claim 1, characterized in that When multiple preparatory words are obtained, the following steps are executed: Determine whether the various preparatory words point to the same meaning, and calculate whether the proportion of the number of preparatory words pointing to the same meaning exceeds the preset value; If the proportion of the number of preparatory words pointing to the same meaning is greater than the preset value, continue to execute the preloading of the corresponding interpretation model in the preset database in advance; If the proportion of the number of preparatory words pointing to the same meaning is less than or equal to the preset value, continuously obtain whether other preparatory words continue to appear within the preset time, and in real time determine whether it exceeds the preset value after the newly added preparatory words are added within the preset time. If it exceeds, continue to execute the preloading of the corresponding interpretation model in the preset database in advance. If it does not exceed, use the filtered preparatory words as the loading words for preloading.
7. An Internet of Things multimedia monitoring and management system for a venue according to claim 5, characterized in that, The specific operation of using the filtered preparatory words as the loading words for preloading is as follows: Use the basic requirement words among the multiple preparatory words as the main loading words, and other words as the parameter words; Select the interpretation model corresponding to the main loading words as the basic interpretation method, and adjust the basic interpretation method in combination with the corresponding interpretation model parameters in the parameter words.
8. The multimedia monitoring and management system for the Internet of Things in a venue according to claim 1, characterized in that, It also includes a derivative recommendation module, and the specific workflow is as follows: When the data analysis module obtains preparatory words from the voice interaction data, automatically generate a set of potential demand words related to the preparatory words; Store the set of potential demand words as a preloading item in the temporary buffer, and start the demand monitoring timer; During the timing of the demand monitoring timer, continuously monitor the voice interaction data of the target population; If an execution word matching any word in the set of potential demand words is detected in the voice interaction data of the target population, call the derivative product interpretation model corresponding to the execution word through the model call module, and fuse the derivative product interpretation model with the current basic interpretation model and recommend it to the target population; If no matching execution word is detected when the demand monitoring timer times out, clear the preloading items in the temporary buffer.
9. The multimedia monitoring and management system for the Internet of Things in a venue according to claim 8, wherein, During the timing of the demand monitoring timer, if a specific behavior combination of the target population is detected, accelerate the trigger of derivative recommendation: when the user stays in front of a certain commodity for more than the preset time and a preparatory word appears in the voice interaction, immediately call the derivative product interpretation model without waiting for the demand monitoring timer to end.
10. The multimedia monitoring and management system for the venue Internet of Things according to claim 1, wherein, It also includes a storage module, which is used to store the preset population classification, basic interpretation model, the corresponding relationship database between the basic interpretation model and the expected behavioral feature data, and the corresponding relationship between the preparatory words and the interpretation model.
Citation Information
Patent Citations
Exhibition hall application method and system based on artificial intelligence
CN117473162A
Method and system for carrying out intelligent explanation on exhibition hall large screen
CN118964601A
Intelligent interactive voice explanation system based on AIGC large model
CN119169994A
Digital exhibition hall recommendation system and method based on artificial intelligence
CN119322969A
Education experience system based on virtual reality technology
CN119694171A