A method and system for generating an exhibition voice drive
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-26
- Publication Date
- 2026-08-11
AI Technical Summary
[0004]现有展陈语音驱动生成技术在处理展示场景交互时,通常仅依赖基础语义解析机制进行内容调度,未深入剖析交互过程中的物理声学特征变化轨迹,难以精确捕捉交互意图的动态切换时机,容易导致展示内容响应与实际交互节奏脱节,同时在将生成内容映射至界面时,往往基于固定空间分配逻辑,未实时监测并结合显示屏幕底层像素级视觉变化状态评估实际空闲程度,使得内容落位极易干扰正在进行的其他展示进程,造成显示空间利用效率低下以及多区域呈现的视觉冲突,严重影响展陈信息组织呈现质量与整体交互控制效果
本发明中,通过提取展陈语音模拟信号中的音频能量及静音时长特征构建序列,依此界定交互阶段变点以拆分待分配内容,有效克服常规手段交互节奏脱节缺陷,实现对交互时机与意图的敏锐捕捉与精准响应,同步采集连续刷新周期内屏幕区域像素色彩值计算色彩变化率,据此精准筛查真实空闲状态的物理像素矩阵区域构建可用显示槽位,依据内容与槽位属性特征制定比例标识并计算分配代价,利用矩阵变换寻优求解显示空间最优落位匹配关系,避免固定分配造成的视觉冲突与空间浪费,确保多区域数字媒体内容动态渲染的协调性与视觉呈现融合。
Smart Images

Figure CN122266359B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of acoustic interaction technology, and in particular to a method and system for generating presentation voice. Background Technology
[0002] The field of acoustic interaction technology involves technologies that utilize sound signals to achieve information transmission and control between people and devices or systems. It covers related content such as voice acquisition, voice recognition, semantic understanding, and voice synthesis. Its main purpose is to realize information input and output through sound as a natural medium. It is widely used in industries such as intelligent exhibitions, digital museums, human-computer interaction terminals, intelligent guides, and multimedia display systems.
[0003] Among them, the exhibition voice-driven generation method refers to a class of methods that generate, schedule or control the exhibition content by collecting and processing user voice commands in the exhibition scene. It usually relies on speech recognition technology, semantic parsing technology and content management and generation mechanism to realize the organization, presentation and interactive control of exhibition information.
[0004] Existing exhibition voice-driven generation technologies typically rely solely on basic semantic parsing mechanisms for content scheduling when handling interactive exhibition scenarios. They fail to deeply analyze the physical acoustic characteristics changes during the interaction process, making it difficult to accurately capture the dynamic switching timing of interactive intentions. This can easily lead to a disconnect between the response of the exhibition content and the actual interaction rhythm. Furthermore, when mapping the generated content to the interface, it often relies on fixed space allocation logic without real-time monitoring and assessment of the actual idle level based on the pixel-level visual changes at the display screen's underlying layer. This makes it easy for content placement to interfere with other ongoing exhibition processes, resulting in low display space utilization efficiency and visual conflicts between multiple areas. This severely impacts the quality of exhibition information organization and presentation, as well as the overall interactive control effect. Summary of the Invention
[0005] The purpose of this invention is to address the shortcomings of existing technologies by proposing a presentation-driven speech generation method and system.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: a presentation speech-driven generation method, comprising the following steps: S1: Obtain the simulated speech signal for the exhibition, determine the waveform sequence of the exhibition speech, determine the audio energy feature value and the silence duration feature value based on the waveform sequence of the exhibition speech, and generate the statistical feature sequence of the exhibition speech. S2: Determine the change points of the exhibition interaction stage based on the exhibition voice statistical feature sequence, and combine the change points of the exhibition interaction stage with the exhibition voice waveform sequence to generate the exhibition content to be assigned sequence. S3: Obtain the regional pixel color values of the exhibition splicing display wall within the continuous refresh cycle, calculate the color change rate of the exhibition interface based on the regional pixel color values, filter the screen space pixel matrix areas that are in an idle state based on the color change rate of the exhibition interface, and determine the set of available display slots for the exhibition based on multiple screen space pixel matrix areas. S4: Based on the sequence of content to be allocated in the exhibition, determine the first ratio identifier; based on the set of available display slots in the exhibition, determine the second ratio identifier; based on the first ratio identifier and the second ratio identifier, allocate the exhibition space to obtain the optimal allocation result of the exhibition space. S5: Obtain the digital media file path data of the exhibition space corresponding to the optimal placement allocation result of the exhibition space, and extract the starting horizontal and vertical coordinates and the ending horizontal and vertical coordinates corresponding to the pixel matrix region of the screen space. Based on the digital media file path data of the exhibition space, the starting horizontal and vertical coordinates and the ending horizontal and vertical coordinates, generate multi-region rendering control instructions for the exhibition space.
[0007] As a further aspect of the present invention, the exhibition voice statistical feature sequence includes audio energy feature value and silence duration feature value; the exhibition content to be allocated sequence includes exhibition interaction stage change points and exhibition voice waveform sequence; the exhibition available display slot set includes screen space pixel matrix region; the optimal placement allocation result of exhibition space includes first ratio identifier, second ratio identifier and exhibition space placement; and the exhibition multi-region rendering control instruction includes exhibition digital media file path data, starting horizontal and vertical coordinates, and ending horizontal and vertical coordinates.
[0008] As a further aspect of the present invention, the step of obtaining the statistical feature sequence of the presentation speech specifically includes: S111: Collect the exhibition scene's audio analog signal during the continuous interaction period with the audience, perform analog-to-digital conversion, obtain the exhibition audio waveform sequence, set a time sliding window, extract the exhibition audio waveform amplitude value of the exhibition audio waveform sequence within the time sliding window, and calculate the total audio energy in the current time period based on the amplitude values of all exhibition audio waveforms within the time sliding window. S112: Extract the time interval between adjacent pronunciation peaks in the corresponding time dimension of the displayed speech waveform sequence, and calculate the average silence duration of all adjacent pronunciation peak intervals; S113: Calculate the proportion of the total audio energy under the preset maximum energy threshold as the audio energy feature value, and at the same time calculate the proportion of the average silence duration under the preset maximum duration threshold as the silence duration feature value. Combine the two feature values at the same time to generate the presentation speech statistical feature sequence.
[0009] As a further aspect of the present invention, the step of obtaining the sequence of content to be assigned for display specifically includes: S211: Extract the audio energy feature value and silence duration feature value from the statistical feature sequence of the exhibition speech, calculate the posterior probability of the audio energy feature value and silence duration feature value in each preset interaction state using the Bayesian online change point detection algorithm, obtain the interaction state attribution probability, and compare the interaction state attribution probability of the same interaction state in adjacent time periods to obtain the interaction state probability change value. S212: Compare the probability change value of the interaction state with the preset stage switching threshold to determine the change point of the exhibition interaction stage, and perform text conversion operation on the exhibition voice waveform sequence according to the change point of the exhibition interaction stage to extract the request string; S213: The request string is split into exhibition object information referring to the entity name of the exhibit and exhibition description information referring to the decorative attributes of the exhibit, and then combined and arranged to obtain the sequence of exhibit content to be allocated.
[0010] As a further aspect of the present invention, the step of obtaining the set of display slots for exhibition is specifically as follows: S311: Monitor the pixel color values of the screen-mapped memory area within the continuous refresh cycle of the exhibition splicing display wall, analyze the pixel color difference of the screen-mapped memory area within adjacent refresh cycles, and obtain the proportion of the pixel color values of the screen-mapped memory area in the initial refresh cycle to obtain the color change rate of the exhibition interface. S312: Extract the continuous period statistics of the color of the display interface when the color change rate of the display interface is zero, filter the physical areas mapped by the continuous period statistics of the color of the display interface being greater than the preset idle period threshold, and obtain the screen space pixel matrix area. S313: Based on the boundary of the screen space pixel matrix region, perform a slot mapping and division operation on the overall display space of the exhibition splicing display wall, summarize the slots of multiple screen space pixel matrix regions, and establish a set of available display slots for the exhibition.
[0011] As a further aspect of the present invention, the step of obtaining the optimal placement result of the exhibition space is specifically as follows: S411: Assign a first proportional identifier representing the content attribute to the display object information and display description information in the display content sequence to be allocated, and assign a second proportional identifier representing the slot function attribute to each screen space pixel matrix area in the set of available display slots, thus forming a display feature proportional identifier group. S412: Extract the first ratio identifier and the second ratio identifier from the exhibition feature ratio identifier group, measure the difference between the first ratio identifier and the second ratio identifier, obtain the exhibition placement allocation cost value, combine multiple exhibition placement allocation cost values, and generate a two-dimensional cost matrix. S413: Perform matrix transformation operations on the two-dimensional value matrix using the Hungarian algorithm to find an independent set of all-zero elements. Based on the independent set of all-zero elements, extract the information of the exhibited objects, the information of the exhibited descriptions, and the pairing row and column relationship data of each screen space pixel matrix region. Assign state values to the pairing row and column relationship data as the optimal placement result of the exhibit space.
[0012] As a further aspect of the present invention, the step of obtaining the multi-region rendering control command for display specifically includes: S511: Extract the exhibition object information and exhibition description information associated with the optimal placement result of the exhibition space, and perform resource file data retrieval from the local repository according to the string matching rules involved in the exhibition object information and exhibition description information to find the digital media file path data of the exhibition. S512: Perform a spatial range positioning and parsing operation on the screen space pixel matrix region pointed to by the optimal placement allocation result of the exhibition space, extract the corresponding screen space start horizontal and vertical coordinates and screen space end horizontal and vertical coordinates according to the screen space pixel matrix region mapping relationship, and generate a region horizontal and vertical coordinate sequence. S513: Extract the screen space start and end coordinates from the horizontal and vertical coordinate sequence of the region, perform low-level data frame splicing encoding on the display digital media file path data, the screen space start and end coordinates, and generate display multi-region rendering control instructions.
[0013] An exhibition voice-driven generation system, the system comprising: The speech feature extraction module acquires the simulated speech signal of the exhibition, determines the waveform sequence of the exhibition speech, determines the audio energy feature value and the silence duration feature value based on the waveform sequence of the exhibition speech, and generates the statistical feature sequence of the exhibition speech. The content to be assigned module determines the change points of the exhibition interaction stage based on the exhibition voice statistical feature sequence, and combines the change points of the exhibition interaction stage with the exhibition voice waveform sequence to generate the exhibition content to be assigned sequence. The available display slot determination module obtains the regional pixel color values of the exhibition splicing display wall within a continuous refresh cycle, calculates the color change rate of the exhibition interface based on the regional pixel color values, filters the screen space pixel matrix areas that are in an idle state based on the color change rate of the exhibition interface, and determines the set of available display slots for the exhibition based on multiple screen space pixel matrix areas. The optimal space allocation module determines a first ratio identifier based on the sequence of exhibits to be allocated, determines a second ratio identifier based on the set of available display slots, and allocates exhibit space based on the first ratio identifier and the second ratio identifier to obtain the optimal space allocation result. The rendering control instruction generation module obtains the digital media file path data of the exhibition space corresponding to the optimal placement allocation result of the exhibition space, and extracts the starting and ending horizontal and vertical coordinates of the pixel matrix region of the screen space. Based on the digital media file path data, the starting and ending horizontal and vertical coordinates, and the ending horizontal and vertical coordinates, it generates multi-region rendering control instructions for the exhibition.
[0014] Compared with the prior art, the advantages and positive effects of the present invention are as follows: In this invention, a sequence is constructed by extracting audio energy and silence duration features from the simulated audio signal of the exhibition. Based on this, the interaction stage change points are defined to split the content to be allocated, effectively overcoming the defect of the interaction rhythm being out of sync with conventional methods. This enables keen capture and precise response to the timing and intent of the interaction. Simultaneously, the color change rate of the pixel values of the screen area within the continuous refresh cycle is calculated. Based on this, the physical pixel matrix area in the real idle state is accurately screened to construct available display slots. The proportion is determined according to the content and slot attribute characteristics, and the allocation cost is calculated. The optimal placement matching relationship of the display space is solved by matrix transformation, avoiding visual conflicts and space waste caused by fixed allocation, and ensuring the coordination and visual integration of dynamic rendering of digital media content in multiple areas. Attached Figure Description
[0015] Figure 1 This is a schematic diagram of the workflow of the present invention; Figure 2 This is a flowchart illustrating the process of obtaining the statistical feature sequence of the presentation speech in this invention; Figure 3 This is a flowchart illustrating the process of obtaining the sequence of content to be assigned for exhibition purposes according to the present invention. Figure 4 This is a flowchart illustrating the process of obtaining a set of available display slots for exhibition purposes according to the present invention. Figure 5 This is a flowchart illustrating the process of obtaining the optimal placement allocation result of the exhibition space according to the present invention; Figure 6 This is a flowchart illustrating the process of obtaining multi-region rendering control instructions for the present invention. Detailed Implementation
[0016] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments; it should be understood that the specific embodiments described herein are merely for explaining the invention and are not intended to limit the invention. Please see Figure 1 This invention provides a technical solution, a method for generating presentation voice, comprising the following steps: S1: Obtain the simulated speech signal for the exhibition, determine the waveform sequence of the exhibition speech, determine the audio energy feature value and the silence duration feature value based on the waveform sequence of the exhibition speech, and generate the statistical feature sequence of the exhibition speech. S2: Determine the variable points of the exhibition interaction stage based on the statistical feature sequence of the exhibition voice, and combine the variable points of the exhibition interaction stage with the exhibition voice waveform sequence to generate the sequence of exhibition content to be assigned. S3: Obtain the regional pixel color values of the exhibition splicing display wall within the continuous refresh cycle, calculate the color change rate of the exhibition interface based on the regional pixel color values, filter the screen space pixel matrix areas that are in an idle state based on the color change rate of the exhibition interface, and determine the set of available display slots for the exhibition based on multiple screen space pixel matrix areas. S4: Determine the first ratio identifier based on the sequence of content to be allocated in the exhibition, determine the second ratio identifier based on the set of available display slots in the exhibition, and allocate the exhibition space according to the first ratio identifier and the second ratio identifier to obtain the optimal allocation result of the exhibition space. S5: Obtain the digital media file path data corresponding to the optimal placement allocation result of the exhibition space, and extract the starting and ending horizontal and vertical coordinates corresponding to the pixel matrix area of the screen space. Based on the digital media file path data, the starting and ending horizontal and vertical coordinates, generate multi-area rendering control instructions for the exhibition.
[0017] The exhibition voice statistical feature sequence includes audio energy feature value and silence duration feature value; the exhibition content to be allocated sequence includes exhibition interaction stage change points and exhibition voice waveform sequence; the exhibition available display slot set includes screen space pixel matrix area; the optimal placement allocation result of exhibition space includes first ratio identifier, second ratio identifier and exhibition space placement; the exhibition multi-area rendering control instructions include exhibition digital media file path data, starting horizontal and vertical coordinates and ending horizontal and vertical coordinates.
[0018] Please see Figure 2 The specific steps for obtaining the statistical feature sequence of the presented speech are as follows: S111: Collect the exhibition scene's audio analog signal during the continuous interaction period with the audience, perform analog-to-digital conversion, obtain the exhibition audio waveform sequence, set a time sliding window, extract the exhibition audio waveform amplitude value of the exhibition audio waveform sequence within the time sliding window, and calculate the total audio energy in the current time period based on the amplitude values of all exhibition audio waveforms within the time sliding window. Acoustic changes are continuously received by sound pickup devices within the exhibition scene during the continuous interaction period of the audience. These acoustic changes are then converted into continuous electrical signals and input into an analog-to-digital converter (ADC). The ADC divides the continuous electrical signal into discrete sampling points arranged in chronological order according to a fixed sampling frequency, and writes the corresponding level value of each sampling point into an exhibition speech waveform sequence. A time sliding window moves sequentially along the exhibition speech waveform sequence from front to back, capturing all sampling points within the current window's coverage area with each movement. Subsequently, the amplitude value of the exhibition speech waveform is read for each sampling point within the current window. When reading the amplitude value of the exhibition speech waveform, the zero point of amplitude is first determined, and then the deviation of the current sampling point from the zero point is taken as the effective amplitude of the current sampling point. When the current sampling point is above the zero point, the positive deviation is recorded directly; when the current sampling point is below the zero point, it is first converted into an absolute deviation before recording to avoid the positive and negative directions canceling each other out. After all the exhibition speech waveform amplitude values within the current window have been read, they are accumulated point by point in the sampling order to obtain the total audio energy for the current time period. After the time sliding window moves to the next position, the same process is repeated until all sampling points within the continuous interaction time period of the audience are processed. If there is an amplitude saturation point in the current window, the saturation point is first marked as the truncated sampling point, and then the average amplitude of the adjacent valid sampling points before and after the truncated sampling point is used to replace the amplitude of the truncated sampling point before being added to the accumulation; if there is no amplitude saturation point in the current window, the accumulation is performed directly; if the number of sampling points in the current window is less than a complete window, the actual number of sampling points is retained and the accumulation is completed according to the actual number of sampling points, and no external data is added.
[0019] S112: Extract the time interval between adjacent pronunciation peaks in the corresponding time dimension of the displayed speech waveform sequence, and calculate the average silence duration of all adjacent pronunciation peak intervals; The amplitude relationship between each sampling point and its predecessor and successor is examined sequentially over time. Positions where the amplitude of the preceding sampling point continuously rises and then declines at the current sampling point are recorded as candidate sound peaks. These candidate sound peaks are then arranged sequentially according to their appearance order. After the candidate sound peaks are arranged, the position difference between consecutive candidate sound peaks on the time axis is read to obtain the interval time between adjacent sound peaks. After extracting the interval time between adjacent sound peaks, it is not directly written into the silence statistics. Instead, it is first checked whether there is a complete trough between adjacent sound peaks that continuously crosses zero and then rises again. When there is a complete trough between adjacent sound peaks, the current interval time is written into the silence duration statistics; when there are only minor local fluctuations between adjacent sound peaks without forming a complete trough, the two candidate sound peaks are merged into one sound peak, and the interval time is not counted separately; when there is a long flat segment between adjacent sound peaks, the current interval time is retained and marked as a long silence interval. After determining the intervals between all adjacent articulation peaks, the intervals are summed sequentially and then divided by the number of adjacent articulation peak intervals to obtain the average silence duration. If only one articulation peak is extracted from the presented speech waveform sequence, there are no adjacent articulation peak intervals, and the average silence duration is recorded as 0. If two or more articulation peaks are extracted, the average is calculated based on the actual number of adjacent articulation peak intervals. If multiple consecutive intervals are consecutive and fall within the same silence segment, each interval is retained separately before being included in the average calculation; they are not merged beforehand.
[0020] S113: Calculate the proportion of the total audio energy under the preset maximum energy threshold as the audio energy feature value, and at the same time calculate the proportion of the average silence duration under the preset maximum duration threshold as the silence duration feature value. Concatenate the two feature values at the same time to generate the presentation speech statistical feature sequence. The preset maximum energy threshold is determined by the historical sequence of the total audio energy under the presented speech input state. First, the total audio energy under the presented speech input state is statistically analyzed, then the average value is taken and superimposed with twice the standard deviation to obtain the preset maximum energy threshold. The preset maximum duration threshold is determined by the historical sequence of the interval time between adjacent sound peaks under interactive pause state and environmental silence state. First, the silence interval time is statistically analyzed, then the average value is taken and superimposed with twice the standard deviation to obtain the preset maximum duration threshold. After comparing the total audio energy with the preset maximum energy threshold, if the total audio energy is greater than the preset maximum energy threshold, the audio energy feature value is recorded as 1; if the total audio energy is equal to the preset maximum energy threshold, the audio energy feature value is also recorded as 1; if the total audio energy is less than the preset maximum energy threshold, the audio energy feature value is converted according to the ratio of the two. After comparing the average silence duration with a preset maximum duration threshold, if the average silence duration is greater than the preset maximum duration threshold, the silence duration feature value is recorded as 1; if the average silence duration is equal to the preset maximum duration threshold, the silence duration feature value is also recorded as 1; if the average silence duration is less than the preset maximum duration threshold, the silence duration feature value is converted according to the ratio of the two. The audio energy feature value and the silence duration feature value at the current time are concatenated end-to-end according to the same time index to form a statistical feature item at a time point. Then, the statistical feature items at all time points are arranged in chronological order to obtain the presentation speech statistical feature sequence. If the current time point lacks the total audio energy, no statistical feature item is generated for the current time point; if the current time point lacks the average silence duration, no statistical feature item is generated for the current time point either; if both features exist simultaneously, the current time point is normally written into the presentation speech statistical feature sequence.
[0021] Please see Figure 3 The specific steps for obtaining the sequence of content to be assigned in the exhibition are as follows: S211: Extract the audio energy feature value and silence duration feature value from the statistical feature sequence of the exhibition speech. Calculate the posterior probability of the audio energy feature value and silence duration feature value under each preset interaction state using the Bayesian online change point detection algorithm to obtain the interaction state attribution probability. Compare the interaction state attribution probability of the same interaction state under adjacent time periods to obtain the interaction state probability change value.
[0022] Interactive states include exhibition voice input state, exhibition interaction pause state, and exhibition environment silent state; The audio energy feature value sequence and the silence duration feature value sequence are extracted separately, and then state distributions are established for the exhibition voice input state, the exhibition interaction pause state, and the exhibition environment silence state, respectively. When establishing the state distributions, firstly, voice segments that have been aligned and marked within the same exhibition scene are collected. Continuously speaking segments are classified as exhibition voice input states, short pauses within voice segments are classified as exhibition interaction pause states, and segments with no audience voices and only ambient noise are classified as exhibition environment silence states. Then, the frequency distributions of the occurrence intervals of audio energy feature values and silence duration feature values are statistically analyzed for each of the three states. After online detection begins, the predicted probability of the audio energy feature value at the current time point is first obtained from the audio energy distribution of each of the three states, and the predicted probability of the silence duration feature value is then obtained from the silence duration distribution of each of the three states. Finally, the posterior probability of the state at the previous time point, the state transition probability, and the two predicted probabilities at the current time point are all fed into the recursive update relationship of the Bayesian online change point detection algorithm.
[0023] The corresponding recurrence relation is written as ; Where t represents the current time point, which is directly given by the time index of the statistical feature sequence of the exhibited speech; This indicates the current state to be determined; only the following conditions are considered. one of the; This represents the state at the previous time point, used to traverse all possible states that could have been entered at the previous time point; The normalized state traversal index is used to include all presentation voice input states, presentation interaction pause states, and presentation environment silence states in the denominator calculation; g represents the presentation voice input state, which comes from the state labels of continuous vocal segments in the alignment markers; r represents the presentation interaction pause state, which comes from the state labels of short pause segments in the alignment markers; q represents the presentation environment silence state, which comes from the state labels of segments without audience vocalization in the alignment markers. This indicates that the current time point belongs to a state. The posterior probability is obtained from the update result at the current time point; This indicates that the state at the previous time point was... The posterior probability, where This represents the interval between adjacent time points, which is obtained from the difference in adjacent time indices in the current sequence; Indicates from the previous state Transition to the current state The state transition probability is obtained by counting the frequency of adjacent state transitions in the aligned and marked state sequence, and then dividing the corresponding transition frequency by the total transition frequency of the previous state. This indicates that the audio energy feature value at the current time point falls into the state. The predicted probability is obtained by using the current audio energy feature value in the state. The relative frequency is obtained from the corresponding frequency distribution; This indicates that the current time point's silence duration feature value falls into the state. The predicted probability is obtained by using the current silence duration feature value in the state. The relative frequency is obtained from the corresponding frequency distribution; This represents the fusion coefficient of two types of features, with a value ranging from 0 to 1. The greater the dispersion of the audio energy feature value among the three states, the better. The larger the value, the greater the dispersion of the silence duration feature value among the three states. The smaller, The variance is obtained by the ratio of the variance between audio energy eigenvalue states to the total variance of the two types of eigenvalues. This means simultaneously iterating through and summing all current candidate states and all states from the previous time point; This represents the posterior probability of each state at the previous time point; Indicates the previous state Transition to the current traversal state The transition probability; This indicates that the audio energy feature value at the current time point falls into the current traversal state. The predicted probability; This indicates that the silence duration feature value at the current time point falls into the current traversal state. The predicted probability; and These represent the fusion weights for the predicted probability of audio energy feature values and the predicted probability of silence duration feature values, respectively.
[0024] For example, first, directly read the posterior probability of the previous time point from the state update result of the previous time point to obtain... Then, from the aligned state sequence, the frequency of adjacent state transitions is counted and converted into transition probabilities to obtain... Next, the audio energy characteristic value and the silence duration characteristic value at the current time point are substituted into the frequency distribution of the three states, respectively, to obtain... as well as Finally, substituting the variance between states of the audio energy characteristic value and the variance between states of the silence duration characteristic value into the coefficients, we obtain the relationship. In the current scenario, the posterior probability at the previous time point is... The state transition probability takes the value of The predicted probability at the current time point is ; The fusion coefficient is set to a value of .
[0025] The numerator for the voice input state of the presentation is written separately as: ;
[0026] Substituting the values from the current scene, we get:
[0027] The molecules in the paused state of the exhibition interaction are written separately as: ; Substituting the values from the current scene, we get:
[0028] The molecule in the silent state of the exhibition environment is written separately as: ; Substituting the values from the current scene, we get:
[0029] The sum of the three state numerators is used as the normalized denominator to obtain the result. Substituting these values into the posterior probability relation, we obtain... The current time point has the highest posterior probability of being in a paused interactive state, so it is prioritized for this state. Then, the posterior probabilities of the same interactive state at the current time point and the previous time point are subtracted, and the absolute value is taken to obtain the change in interactive state probability. When the posterior probability of an interactive state is the highest, the current time point is prioritized for that state. If the posterior probabilities of two interactive states are equal, the posterior probabilities of the two corresponding interactive states at the previous time point are compared, and the state with the higher posterior probability at the previous time point is retained. When the posterior probabilities of all three interactive states are 0, the current time point retains the state label from the previous time point, and the change in interactive state probability at the current time point is recorded as 0.
[0030] S212: Compare the probability change value of the interactive state with the preset stage switching threshold to determine the change point of the exhibition interaction stage, and perform text conversion operation on the exhibition voice waveform sequence according to the change point of the exhibition interaction stage to extract the request string; The probability changes of the exhibition's voice input state, interactive pause state, and silent environment state at adjacent time points are arranged chronologically and then merged into a stage switching reference sequence. The preset stage switching threshold is determined using this reference sequence by first averaging all changes and then adding twice the standard deviation. When comparing the probability change of the same interactive state between the current and previous time points with the preset stage switching threshold, if at least one interactive state's probability change is greater than the threshold, the current time boundary is recorded as a stage change point; if at least one interactive state's probability change is equal to the threshold, the current time boundary is also recorded as a stage change point; if all interactive state probabilities are less than the threshold, the current time boundary is not recorded as a stage change point. After determining the stage change points, the exhibition's voice waveform sequence is first segmented into continuous segments according to the change point positions, and then the complete waveform data is read from the segment containing the current request for text conversion. During text conversion, the current segment is first divided into short speech frames. Then, speech descriptive quantities such as pronunciation fluctuations, formant trends, voiced / voiced transition positions, and pause positions are extracted from each speech frame. Subsequently, the descriptive quantities of consecutive speech frames are matched segment by segment against the exhibit name vocabulary, descriptive attribute vocabulary, and action trigger vocabulary in the exhibition scene. Successfully matched segments are concatenated into a request string according to their chronological order. If only exhibit name segments are matched in the current segment but not descriptive attribute segments, the request string only retains exhibit name segments. If only descriptive attribute segments are matched in the current segment but not exhibit name segments, the request string retains descriptive attribute segments and awaits completion in subsequent segmentation stages. If both exhibit name segments and descriptive attribute segments are successfully matched in the current segment, the complete request string is concatenated according to the original speech order.
[0031] S213: Split the request string into exhibition object information that refers to the name of the exhibit entity and exhibition description information that refers to the decorative attributes of the exhibit, and combine and arrange them to obtain the sequence of exhibit content to be allocated; The request string is segmented into consecutive segments based on semantic boundaries. Each segment is then checked for its correspondence with the exhibit entity name table and exhibit decoration attribute table. If the current segment completely matches an entity name in the exhibit entity name table, it is recorded as exhibit object information. If the current segment completely matches an attribute name in the exhibit decoration attribute table, it is recorded as exhibit description information. If the current segment matches both the exhibit entity name table and the exhibit decoration attribute table, the preceding and following segments in the request string are checked first. If the preceding and following segments contain exhibit names, it is treated as exhibit description information; if the preceding and following segments contain description attributes, it is treated as exhibit object information. If the current segment does not match in either table, it is removed from the segmentation results and not included in the content to be assigned. After the exhibit object information and exhibit description information are determined, they are recombined according to their order of appearance in the request string. When a single display object is followed by multiple consecutive display descriptions, these descriptions are combined with the same display object sequentially to obtain multiple content items to be assigned. When multiple display objects are followed by only one description, the description is assigned to each object based on proximity, resulting in multiple content items to be assigned. When a description precedes a display object, the nearest object is searched for and combined. When duplicate descriptions exist in the request string, only the first occurrence of the description with the same name is retained for combination. After combination, the results are arranged sequentially to obtain the sequence of display content to be assigned.
[0032] Please see Figure 4 The specific steps for obtaining the set of available display slots for the exhibition are as follows: S311: Monitor the pixel color values of the screen-mapped memory area within the continuous refresh cycle of the exhibition splicing display wall, analyze the pixel color difference of the screen-mapped memory area within adjacent refresh cycles, and obtain the proportion of the pixel color values of the screen-mapped memory area in the initial refresh cycle to obtain the color change rate of the exhibition interface. Within a continuous refresh cycle, the red, green, and blue channel values of each pixel are read cycle by cycle. The absolute values of the three channel differences for the same pixel between the current and previous refresh cycles are then summed to obtain the single-pixel color difference. After determining the single-pixel color difference for each pixel, all single-pixel color differences are accumulated according to the pixel's spatial location to obtain the total pixel color difference for the screen-mapped video memory area in the current refresh cycle. When determining the initial refresh cycle, the first refresh cycle without content switching is selected after a stable display of the presentation interface. The three channel values of all pixels in the reference cycle are then summed to obtain the total pixel color difference for the initial refresh cycle. The color change rate of the presentation interface is obtained by the ratio of the total pixel color difference for the current refresh cycle to the total pixel color difference for the initial refresh cycle. When the total pixel color value in the initial refresh cycle is greater than 0, the color change rate of the display interface is directly calculated based on the ratio. When the total pixel color value in the initial refresh cycle is equal to 0, the first refresh cycle with a total pixel color value greater than 0 is selected as the replacement reference cycle, and the ratio calculation is performed again. When the total pixel color difference in the current refresh cycle is equal to 0, the color change rate of the display interface in the current refresh cycle is recorded as 0. After all refresh cycles are processed, a sequence of display interface color change rates arranged by time is obtained.
[0033] S312: Extract the continuous period statistics of the color of the display interface when the color change rate is zero, filter the physical areas mapped by the continuous period statistics of the color of the display interface that are greater than the preset idle period threshold, and obtain the screen space pixel matrix area. The system checks the color change rate of the display interface for each refresh cycle to ensure it is zero. The number of consecutive refresh cycles with a zero change rate is then used as the display interface color continuous cycle statistic. The preset idle cycle threshold is determined by the statistical results of consecutive zero-value cycles in historical idle interface segments. First, the average of the historical consecutive zero-value cycle statistics is taken, then multiplied by one standard deviation to obtain the preset idle cycle threshold. The current area's display interface color continuous cycle statistic is compared with the preset idle cycle threshold. If the display interface color continuous cycle statistic is greater than the preset idle cycle threshold, the physical display position corresponding to the current area is selected into the screen space pixel matrix region; if the display interface color continuous cycle statistic is equal to the preset idle cycle threshold, the current area is also selected into the screen space pixel matrix region; if the display interface color continuous cycle statistic is less than the preset idle cycle threshold, the current area is not selected into the screen space pixel matrix region. During the continuous zero-value cycle statistics, if the color change rate of the display interface is greater than 0 in any one of the intermediate refresh cycles, the continuous count is immediately cleared and restarted from the next refresh cycle; if the color change rate of the display interface is equal to 0 for multiple consecutive refresh cycles, the continuous count continues to accumulate; if the continuous count is not interrupted when the refresh cycle sequence ends, the screen space pixel matrix area corresponding to the current continuous count is directly output according to the end position.
[0034] S313: Based on the boundaries of the screen space pixel matrix region, perform slot mapping and division operations on the overall display space of the exhibition splicing display wall, summarize the slots of multiple screen space pixel matrix regions, and establish a set of available display slots for the exhibition. The boundary pixel coordinates of each pixel matrix region are read and then projected onto the overall display space of the exhibition wall to obtain the corresponding space occupancy range. After the space occupancy range is determined, region segmentation is performed along the horizontal and vertical boundaries of the overall display space to form display partitions consistent with the boundaries of the pixel matrix regions. When two pixel matrix regions are connected end-to-end on the horizontal boundary and their vertical coverage overlaps, the two pixel matrix regions are merged into the same display slot; when two pixel matrix regions are connected end-to-end on the vertical boundary and their horizontal coverage overlaps, the two pixel matrix regions are merged into the same display slot; when there are intervening pixel bands between two pixel matrix regions, they are each retained as independent display slots; when two pixel matrix regions overlap, the smallest bounding rectangle of the overlapping part is first taken, and then the boundaries of the two regions are re-corrected using the smallest bounding rectangle. After all region boundaries are processed, the starting boundary, ending boundary, and slot number of each segmentation result are written, and then all slots are summarized to establish a set of available display slots for the exhibition. When the partitioning result forms only one independent region, only one slot is retained in the set of available display slots; when the partitioning result forms multiple independent regions, multiple slots are retained in the set of available display slots in spatial order; when none of the pixel matrix regions meet the independent partitioning condition, the set of available display slots is recorded as an empty set.
[0035] Please see Figure 5 The specific steps for obtaining the optimal placement result of the exhibition space are as follows: S411: Assign a first proportional identifier representing the content attribute to the display object information and display description information in the display content sequence to be allocated, and assign a second proportional identifier representing the slot function attribute to each screen space pixel matrix area in the set of available display slots, thus forming a display feature proportional identifier group; For each content item to be assigned, the character lengths of the display object information and display description information are extracted. These two character lengths are then normalized to generate a first proportional identifier representing the content attribute. When both the display object information and display description information character lengths are greater than 0, the first proportional identifier is generated directly based on their ratio to the total length. When the display object information character length is greater than 0 and the display description information character length is equal to 0, the first proportional identifier is recorded as object information percentage 1 and description information percentage 0. When the display object information character length is equal to 0 and the display description information character length is greater than 0, the first proportional identifier is recorded as object information percentage 0 and description information percentage 1. Subsequently, for each screen space pixel matrix region, the region width and region height are read, and a second proportional identifier representing the slot's functional attribute is generated based on the area of the reserved title area and the area of the description area. The division of the title area and description area does not depend on external additional content; instead, based on the slot boundary, an object information display band is drawn at the top, and a description information display band is drawn in the remaining area, and then the areas of the two parts are calculated separately. After the first ratio identifiers of all content items to be assigned and the second ratio identifiers of all slots are generated, they are combined according to the order of content items and slots to form a display feature ratio identifier group.
[0036] S412: Extract the first and second proportion identifiers from the exhibition feature proportion identifier group, measure the difference between the first and second proportion identifiers, obtain the exhibition placement allocation cost value, combine multiple exhibition placement allocation cost values, and generate a two-dimensional cost matrix. The first proportion identifier is extracted from each content item, and the second proportion identifier is extracted from each slot. Then, each item is compared along its corresponding dimension. The display placement allocation cost is calculated using the difference in the proportion of object information and the difference in the proportion of explanatory information. First, the absolute difference between the proportion of object information and the proportion of object information carrying capacity is taken, then the absolute difference between the proportion of explanatory information and the proportion of explanatory information carrying capacity is taken. Finally, the two absolute differences are summed to obtain the display placement allocation cost. When the first and second proportion identifiers are completely equal in both proportions, the display placement allocation cost is recorded as 0; when only one of the two proportions is equal, the display placement allocation cost consists only of the absolute differences of the unequal items; when neither proportion is equal, the display placement allocation cost is the sum of the two absolute differences. Each content item is compared with each slot in the same way, resulting in multiple sets of display placement allocation cost values. After generating multiple sets of display placement allocation cost values, they are written in rows according to the content item arrangement direction and in columns according to the slot arrangement direction, generating a two-dimensional cost matrix. When the number of content items is greater than the number of slots, the two-dimensional value matrix is retained as a non-square matrix according to the actual number of rows and columns; when the number of content items is equal to the number of slots, the two-dimensional value matrix is written as a square matrix; when the number of slots is greater than the number of content items, the columns corresponding to the extra slots are retained as usual, waiting for the subsequent pairing process to select the best option.
[0037] S413: Perform matrix transformation operations on the two-dimensional value matrix using the Hungarian algorithm to find the independent set of all-zero elements. Based on the independent set of all-zero elements, extract the information of the exhibited objects, the information of the exhibited descriptions, and the paired row and column relationship data of each screen space pixel matrix region. Assign state values to the paired row and column relationship data as the optimal placement result of the exhibit space. Find the minimum value for each row, and then subtract the minimum value for the current row from the minimum value for each row, thus completing the first row transformation. After the row transformation is completed, find the minimum value for each column, and then subtract the minimum value for the current column from the minimum value for each column, thus completing the first column transformation. Then, send the matrix after the row and column transformations into the assignment solution relation.
[0038] The optimal allocation objective is written as: ; It also satisfies the constraints that each row can only select one slot and each column can only be assigned to one content item.
[0039] Row-column relationships are written as ,in .
[0040] Where i represents the content item row index, which is numbered sequentially according to the order of the content items in the sequence of content to be allocated in the display; j represents the slot column index, which is numbered sequentially according to the order of the slots in the set of available display slots in the display. The value in the i-th row and j-th column of the original two-dimensional value matrix is obtained by summing the difference between the proportion of object information and the proportion of explanatory information in the 11th paragraph. Represents the minimum cost value of the i-th row, obtained by... Obtained by comparing each column in the row; This represents the minimum cost value in the j-th column after row transformation, which is obtained by comparing each row in the j-th column. The row and column reduction result of the i-th row and j-th column is obtained by subtracting the minimum value of the corresponding row and the minimum value of the corresponding column from the original value; The indicator value represents whether the i-th content item is assigned to the j-th slot. It is 1 when the assignment is successful and 0 when it is unsuccessful. It is derived from the subsequent independent pairing results. z represents the total cost corresponding to all assignment relationships, which is obtained by summing the original costs of all selected pairing positions.
[0041] For example, the first row contains 0.02, 0.08, and 0.10, the second row contains 0.06, 0.02, and 0.14, and the third row contains 0.10, 0.15, and 0.00. Therefore, the minimum cost in the first row is written as... The minimum cost in line 2 is written as: The minimum cost in line 3 is written as: After substituting the minimum values for the three rows back into the row and column reduction relationships, the row transformation results are as follows: Row 1: 0.00, 0.06, 0.08; Row 2: 0.04, 0.00, 0.12; Row 3: 0.10, 0.15, 0.00. Then, finding the minimum value column by column, the minimum value for the first column is written as follows: The minimum cost in column 2 is written as: The minimum cost in column 3 is written as: After substituting the minimum cost values of the three columns back into the row and column reduction relations, the reduction matrix remains unchanged. Therefore... At this point, three non-conflicting zero-value positions appear in row 1, column 1, row 2, column 2, and row 3, column 3, indicating that each of the three content items can find a unique corresponding slot. Therefore, the indicator value is set to... ,the remaining Set all values to 0. Then substitute the above indicator values into the optimal allocation objective to obtain... Continue to substitute After all other positions are 0, write it as The total cost of 0.04 corresponds to the optimal and unique allocation result under the current matrix, that is, the first content item corresponds to the first slot, the second content item corresponds to the second slot, and the third content item corresponds to the third slot. When a non-conflicting minimum cost pair that can cover all content items already exists in the transformed matrix, the pairing extraction is directly performed; when the non-conflicting minimum cost pairings in the transformed matrix are insufficient to cover all content items, the covering and adjustment process continues until all row and column constraints are satisfied. After the independent pairing is determined, the row and column numbers in the independent pairing are read into the pairing row and column relationship data, and the successfully paired positions are assigned the value of "allocated", while the unpaired positions are assigned the value of "unallocated", thus obtaining the optimal placement allocation result of the display space.
[0042] Please see Figure 6 The specific steps for obtaining multi-region rendering control commands are as follows: S511: Extract the exhibition object information and exhibition description information associated with the optimal placement result of the exhibition space, and perform resource file data retrieval from the local repository based on the string matching rules involved in the exhibition object information and exhibition description information to find the digital media file path data of the exhibition. The system reads the corresponding display object information and display description information item by item according to the assigned positions. Then, it concatenates the display object information and display description information into a search string in a fixed order. The first half of the search string is written into the display object information, and the second half into the display description information. After the search string is generated, it performs full match, prefix match, and segment match sequentially according to the string matching rules in the local repository. When both the display object information and the display description information can be fully matched at the same path level, the corresponding resource file path data is directly output. When the display object information can be fully matched but the display description information cannot, prefix match is performed in the directory corresponding to the display object information, and paths with matching prefixes are retained. When neither the display object information nor the display description information can be fully matched, the search string is further split according to word boundaries, and segment match is performed on each segment. The segments are then sorted from most to least matched, and the path with the highest matching degree is taken as the display digital media file path data. When no full match, prefix match, or segment match is found, no resource file path data is output for the currently assigned position. After string matching is completed, the display digital media file path data corresponding to each assigned position is arranged in pairing order and used as input for subsequent area encoding.
[0043] S512: Perform spatial range positioning and parsing operation on the screen space pixel matrix region pointed to by the optimal placement allocation result of the exhibition space, extract the corresponding screen space start horizontal and vertical coordinates and screen space end horizontal and vertical coordinates according to the screen space pixel matrix region mapping relationship, and generate a region horizontal and vertical coordinate sequence. The boundary positions of the corresponding pixel matrix regions in the screen space are read according to the slot numbers in the optimal placement allocation results of the exhibition space. Then, the starting x-coordinate, starting y-coordinate, ending x-coordinate, and ending y-coordinate are extracted from the mapping relationship. During boundary position extraction, the leftmost pixel in the horizontal direction is written as the starting x-coordinate of the screen space, the topmost pixel in the vertical direction is written as the starting y-coordinate, the rightmost pixel in the horizontal direction is written as the ending x-coordinate, and the bottommost pixel in the vertical direction is written as the ending y-coordinate. When the boundary region is a regular rectangle, the region's x- and y-coordinate sequences are directly generated based on the four boundary values. When the boundary region is an irregular connected region, the smallest bounding rectangle of the irregular connected region is first taken, and then the region's x- and y-coordinate sequences are generated based on the four boundary values of the smallest bounding rectangle. When the boundary region has multiple separate sub-regions, the smallest bounding rectangle of each sub-region is first extracted, and then multiple region x- and y-coordinate sequences are generated sequentially according to the order in which the sub-regions appear. After the region x- and y-coordinate sequences are generated, they are written one-to-one according to the resource path order under the same allocated position to form the coordinate input sequence required for subsequent encoding.
[0044] S513: Extract the screen space start and end coordinates from the region's horizontal and vertical coordinate sequence, perform low-level data frame splicing encoding on the digital media file path data, the screen space start and end coordinates, and generate multi-region rendering control instructions for the presentation. The encoded content is organized line by line, with each resource path corresponding to a set of region coordinates. The resource path field, starting x-coordinate field, starting y-coordinate field, ending x-coordinate field, and ending y-coordinate field are then sequentially concatenated to form a single region control frame. The encoding order of a single region control frame is fixed: the resource file path data is written at the beginning, the screen space starting and ending x- and y-coordinates are written in the middle, and a frame end marker is written at the end. A single region control frame is output completely when both the resource path data and the region x- and y-coordinate sequences exist; if the resource path data exists but the region x- and y-coordinate sequences do not exist, the current resource path data is not encoded; similarly, if the region x- and y-coordinate sequences exist but the resource path data does not exist, the current coordinate sequence is also not encoded. After all single region control frames are generated, they are sequentially concatenated according to the optimal placement allocation result in the presentation space to obtain the multi-region rendering control instructions for the presentation. When multiple single-region control frames exist, they are concatenated sequentially to generate a multi-region rendering control instruction; when only one single-region control frame exists, the single-region control frame is directly used as the multi-region rendering control instruction; when no complete single-region control frames exist, the multi-region rendering control instruction remains empty.
[0045] The entire process is as follows: After the audience member says "Show the description of the patterns on the bronze tripod," the speech is first recorded as a waveform data. Then, the audio energy and pause duration are extracted from the waveform to form a continuously analyzable speech feature sequence. Next, based on these features, it is determined whether the audience member is speaking, pausing, or silent. The actual request speech is then extracted, converted into a request string, and then "bronze tripod" and "description of patterns" are extracted to form the exhibit content to be displayed.
[0046] Meanwhile, the display wall continuously monitors changes in the images of each area, identifying areas that remain unchanged for extended periods as idle areas, and then organizing these idle areas into available display slots. Subsequently, the content "Bronze Tripod, Description of Decoration" is compared with each available slot to select the most suitable display location.
[0047] Once the display location is determined, locate the corresponding image or video file for the "Explanation of Bronze Tripod Decoration" in the local resources and extract the area coordinates corresponding to that display location. Finally, generate rendering control instructions by combining the resource path and area coordinates to display the description of the bronze tripod decoration in the corresponding area of the display wall.
[0048] An exhibition voice-driven generation system, the system comprising: The speech feature extraction module acquires the simulated speech signal of the exhibition, determines the waveform sequence of the exhibition speech, determines the audio energy feature value and the silence duration feature value based on the waveform sequence of the exhibition speech, and generates the statistical feature sequence of the exhibition speech. The content to be assigned module determines the variable points of the exhibition interaction stage based on the statistical feature sequence of the exhibition voice, and combines the variable points of the exhibition interaction stage with the exhibition voice waveform sequence to generate the sequence of content to be assigned in the exhibition. The available display slot determination module obtains the regional pixel color values of the exhibition splicing display wall within a continuous refresh cycle, calculates the color change rate of the exhibition interface based on the regional pixel color values, filters the screen space pixel matrix areas that are in an idle state based on the color change rate of the exhibition interface, and determines the set of available display slots for the exhibition based on multiple screen space pixel matrix areas. The optimal space allocation module determines a first ratio identifier based on the sequence of content to be allocated in the exhibition, determines a second ratio identifier based on the set of available display slots in the exhibition, and allocates exhibition space according to the first ratio identifier and the second ratio identifier to obtain the optimal space allocation result. The rendering control instruction generation module obtains the digital media file path data of the exhibition space corresponding to the optimal placement allocation result of the exhibition space, and extracts the starting and ending horizontal and vertical coordinates of the pixel matrix region of the screen space. Based on the digital media file path data, the starting and ending horizontal and vertical coordinates, the module generates rendering control instructions for multiple areas of the exhibition space.
[0049] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A method for generating presentation speech, characterized in that, Includes the following steps: S1: Obtain the simulated speech signal for the exhibition, determine the waveform sequence of the exhibition speech, determine the audio energy feature value and the silence duration feature value based on the waveform sequence of the exhibition speech, and generate the statistical feature sequence of the exhibition speech. S2: Determine the change points of the exhibition interaction stage based on the exhibition voice statistical feature sequence, and combine the change points of the exhibition interaction stage with the exhibition voice waveform sequence to generate the exhibition content to be assigned sequence. S3: Obtain the regional pixel color values of the exhibition splicing display wall within the continuous refresh cycle, calculate the color change rate of the exhibition interface based on the regional pixel color values, filter the screen space pixel matrix areas that are in an idle state based on the color change rate of the exhibition interface, and determine the set of available display slots for the exhibition based on multiple screen space pixel matrix areas. S4: Based on the sequence of content to be allocated in the exhibition, determine the first ratio identifier; based on the set of available display slots in the exhibition, determine the second ratio identifier; based on the first ratio identifier and the second ratio identifier, allocate the exhibition space to obtain the optimal allocation result of the exhibition space. S5: Obtain the digital media file path data of the exhibition space corresponding to the optimal placement allocation result of the exhibition space, and extract the starting horizontal and vertical coordinates and the ending horizontal and vertical coordinates corresponding to the pixel matrix area of the screen space. Based on the digital media file path data of the exhibition space, the starting horizontal and vertical coordinates and the ending horizontal and vertical coordinates, generate multi-area rendering control instructions for the exhibition space. The specific steps for obtaining the optimal placement result of the exhibition space are as follows: S411: Assign a first proportional identifier representing the content attribute to the display object information and display description information in the display content sequence to be allocated, and assign a second proportional identifier representing the slot function attribute to each screen space pixel matrix area in the set of available display slots, thus forming a display feature proportional identifier group. S412: Extract the first ratio identifier and the second ratio identifier from the exhibition feature ratio identifier group, measure the difference between the first ratio identifier and the second ratio identifier, obtain the exhibition placement allocation cost value, combine multiple exhibition placement allocation cost values, and generate a two-dimensional cost matrix. S413: Perform matrix transformation operations on the two-dimensional value matrix using the Hungarian algorithm to find an independent set of all-zero elements. Based on the independent set of all-zero elements, extract the information of the exhibited objects, the information of the exhibited descriptions, and the pairing row and column relationship data of each screen space pixel matrix region. Assign state values to the pairing row and column relationship data as the optimal placement result of the exhibit space.
2. The exhibition voice-driven generation method according to claim 1, characterized in that, The specific steps for obtaining the statistical feature sequence of the exhibition speech are as follows: S111: Collect the exhibition scene's audio analog signal during the continuous interaction period with the audience, perform analog-to-digital conversion, obtain the exhibition audio waveform sequence, set a time sliding window, extract the exhibition audio waveform amplitude value of the exhibition audio waveform sequence within the time sliding window, and calculate the total audio energy in the current time period based on the amplitude values of all exhibition audio waveforms within the time sliding window. S112: Extract the time interval between adjacent pronunciation peaks in the corresponding time dimension of the displayed speech waveform sequence, and calculate the average silence duration of all adjacent pronunciation peak intervals; S113: Calculate the proportion of the total audio energy under the preset maximum energy threshold as the audio energy feature value, and at the same time calculate the proportion of the average silence duration under the preset maximum duration threshold as the silence duration feature value. Combine the two feature values at the same time to generate the presentation speech statistical feature sequence.
3. The exhibition voice-driven generation method according to claim 2, characterized in that, The specific steps for obtaining the sequence of exhibition content to be assigned are as follows: S211: Extract the audio energy feature value and silence duration feature value from the statistical feature sequence of the exhibition speech, calculate the posterior probability of the audio energy feature value and silence duration feature value in each preset interaction state using the Bayesian online change point detection algorithm, obtain the interaction state attribution probability, and compare the interaction state attribution probability of the same interaction state in adjacent time periods to obtain the interaction state probability change value. S212: Compare the probability change value of the interaction state with the preset stage switching threshold to determine the change point of the exhibition interaction stage, and perform text conversion operation on the exhibition voice waveform sequence according to the change point of the exhibition interaction stage to extract the request string; S213: The request string is split into exhibition object information referring to the entity name of the exhibit and exhibition description information referring to the decorative attributes of the exhibit, and then combined and arranged to obtain the sequence of exhibit content to be allocated.
4. The exhibition voice-driven generation method according to claim 3, characterized in that, The specific steps for obtaining the set of display slots for the exhibition are as follows: S311: Monitor the pixel color values of the screen-mapped memory area within the continuous refresh cycle of the exhibition splicing display wall, analyze the pixel color difference of the screen-mapped memory area within adjacent refresh cycles, and obtain the proportion of the pixel color values of the screen-mapped memory area in the initial refresh cycle to obtain the color change rate of the exhibition interface. S312: Extract the continuous period statistics of the color of the display interface when the color change rate of the display interface is zero, filter the physical areas mapped by the continuous period statistics of the color of the display interface being greater than the preset idle period threshold, and obtain the screen space pixel matrix area. S313: Based on the boundary of the screen space pixel matrix region, perform a slot mapping and division operation on the overall display space of the exhibition splicing display wall, summarize the slots of multiple screen space pixel matrix regions, and establish a set of available display slots for the exhibition.
5. The exhibition voice-driven generation method according to claim 4, characterized in that, The specific steps for obtaining the multi-region rendering control commands for the exhibition are as follows: S511: Extract the exhibition object information and exhibition description information associated with the optimal placement result of the exhibition space, and perform resource file data retrieval from the local repository according to the string matching rules involved in the exhibition object information and exhibition description information to find the digital media file path data of the exhibition. S512: Perform a spatial range positioning and parsing operation on the screen space pixel matrix region pointed to by the optimal placement allocation result of the exhibition space, extract the corresponding screen space start horizontal and vertical coordinates and screen space end horizontal and vertical coordinates according to the screen space pixel matrix region mapping relationship, and generate a region horizontal and vertical coordinate sequence. S513: Extract the screen space start and end coordinates from the horizontal and vertical coordinate sequence of the region, perform low-level data frame splicing encoding on the display digital media file path data, the screen space start and end coordinates, and generate display multi-region rendering control instructions.
6. The exhibition voice-driven generation method according to claim 3, characterized in that, The interactive states include the exhibition voice input state, the exhibition interaction pause state, and the exhibition environment silent state.
7. The exhibition voice-driven generation method according to claim 3, characterized in that, The formula for determining the probability of belonging to an interaction state is as follows: ; in, Indicates the current time point, This indicates the current state to be determined. Indicates the state at the previous point in time. Indicates the index of the state traversal. This indicates that the current time point belongs to a state. The posterior probability is used as the probability of belonging to the current interaction state. Indicates the interval between adjacent time points. Indicates from the previous state Transition to the current state The state transition probability, This indicates that the audio energy feature value at the current time point falls into the state. The predicted probability, This indicates that the current time point's silence duration feature value falls into the state. Predicted probability This means summing the sums of all current candidate states and all states from the previous time point. This represents the posterior probability of each state at the previous time point. Indicates the previous state Transition to the current traversal state The transition probability, This indicates that the audio energy feature value at the current time point falls into the current traversal state. The predicted probability, This indicates that the silence duration feature value at the current time point falls into the current traversal state. The predicted probability, and These represent the fusion weights for the predicted probability of audio energy feature values and the predicted probability of silence duration feature values, respectively.
8. A speech-driven generation system for exhibitions, characterized in that, The exhibition voice-driven generation method according to any one of claims 1-7, wherein the system comprises: The speech feature extraction module acquires the simulated speech signal of the exhibition, determines the waveform sequence of the exhibition speech, determines the audio energy feature value and the silence duration feature value based on the waveform sequence of the exhibition speech, and generates the statistical feature sequence of the exhibition speech. The content to be assigned module determines the change points of the exhibition interaction stage based on the exhibition voice statistical feature sequence, and combines the change points of the exhibition interaction stage with the exhibition voice waveform sequence to generate the exhibition content to be assigned sequence. The available display slot determination module obtains the regional pixel color values of the exhibition splicing display wall within a continuous refresh cycle, calculates the color change rate of the exhibition interface based on the regional pixel color values, filters the screen space pixel matrix areas that are in an idle state based on the color change rate of the exhibition interface, and determines the set of available display slots for the exhibition based on multiple screen space pixel matrix areas. The optimal space allocation module determines a first ratio identifier based on the sequence of exhibits to be allocated, determines a second ratio identifier based on the set of available display slots, and allocates exhibit space based on the first ratio identifier and the second ratio identifier to obtain the optimal space allocation result. The rendering control instruction generation module obtains the digital media file path data of the exhibition space corresponding to the optimal placement allocation result of the exhibition space, and extracts the starting and ending horizontal and vertical coordinates of the pixel matrix region of the screen space. Based on the digital media file path data of the exhibition space, the starting and ending horizontal and vertical coordinates, and the ending horizontal and vertical coordinates, it generates multi-region rendering control instructions for the exhibition. The specific steps for obtaining the optimal placement result of the exhibition space are as follows: S411: Assign a first proportional identifier representing the content attribute to the display object information and display description information in the display content sequence to be allocated, and assign a second proportional identifier representing the slot function attribute to each screen space pixel matrix area in the set of available display slots, thus forming a display feature proportional identifier group. S412: Extract the first ratio identifier and the second ratio identifier from the exhibition feature ratio identifier group, measure the difference between the first ratio identifier and the second ratio identifier, obtain the exhibition placement allocation cost value, combine multiple exhibition placement allocation cost values, and generate a two-dimensional cost matrix. S413: Perform matrix transformation operations on the two-dimensional value matrix using the Hungarian algorithm to find an independent set of all-zero elements. Based on the independent set of all-zero elements, extract the information of the exhibited objects, the information of the exhibited descriptions, and the pairing row and column relationship data of each screen space pixel matrix region. Assign state values to the pairing row and column relationship data as the optimal placement result of the exhibit space.
Citation Information
Patent Citations
Audio endpoint detection and noise reduction method based on power intranet
CN110910906A
Image display system and method of intelligent voice interactive electronic ink screen
CN121171180A