05 / Applied AI / 2026
Speech modelling and detection.
Speech technology / VAD + GMM
01 / SCOPE
Problem and
scope.
The evaluation compares neural voice-activity detectors and classical models across speech corpora, with the cross-corpus result as the main test.
Individual modelling and evaluation work using supplied datasets and practical frontends. The neural VAD and GMM implementations are available in the public repository.
02 / IMPLEMENTATION
Implementation
details.
- 01Acoustic features
- 02Context / sequence
- 03Neural VAD
- 04Post-processing
- 05Cross-corpus test
Compare model families.
Move from a contextual MLP to bidirectional recurrent models and dilated temporal convolutions, examining how temporal context changes detection.
Test on a second corpus.
Use detection-error trade-off curves and equal error rate to compare the selected systems on a different speech corpus.
Compare against GMM baselines.
Use diagonal-covariance Gaussian mixture models for vowels and speaker identification, comparing mixture complexity, utterance aggregation and UBM-MAP adaptation.
03 / RESULTS
Results.
Selected reported cross-corpus results over 495,486 test frames. The tuned TCN includes median post-processing. These figures are retained evaluation results, not a new benchmark run.
Context & limitations
Datasets, labels and checkpoints are excluded from the public repository. The GMM workflow needs a supplied or replacement MFCC frontend. Closed-set speaker results do not establish general-world identification performance.
04 / SUMMARY
The cross-corpus evaluation shows the effect of model family and temporal context; it is not a general speech-recognition claim.Open repository
Productiv / Genesys.
Team-built Rails productivity platform