Changelog
All notable changes to ProteoPy will be documented in this file. The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
[Unreleased]
Added
Preprocessing (
pr.pp):summarize_peptides_by_neighbourhood_union()collapses peptides that overlap in the protein sequence, keeping the most abundant member of each group. Peptide positions are resolved from a FASTA in the same call. A reimplementation of CCprofiler’ssummarizeAlternativePeptideSequences(topN = 1).groups by positional overlap rather than substring containment, so it sees peptide pairs that overlap without either containing the other, and needs no separate modification-summarisation step
selects the most abundant member instead of aggregating;
top_nsums the leading members instead, withkeep_lesscontrolling undersized groupsmissing values are deprioritised rather than removed: an incomplete peptide sorts last and loses to any complete competitor, but survives if its group has no complete member
equal totals are resolved by
tie_break_key, so the result does not depend on input row order; the default sorts non-letters after letters, favouring the unmodified form of an identifieron_unknown_proteinandon_unlocated_peptidedecide whether an unresolvable position raises, skips the peptide, or leaves the position undefined.varis reduced to the peptide-level proteodata columns plus the function’s own output, since the surviving row’s annotations describe one member rather than the group;keep_var_colscarries chosen columns through, aggregated across the group
Fixed
Datasets (
pr.datasets) and Download (pr.download):williams_2018()no longer conflates measured zeros with missing values. Three defects, each masked by another:zeros in
.Xwere coerced tonp.nan, discarding 13,547 genuine measurementscharge-state summation used pandas’ default
min_count=0, so a group with no measurements at all summed to0.0, inventing 3,324the same summation skipped
NaNinside a partially measured group, reporting a partial total as complete for 260 cellsthis changes
.X, and therefore the output ofpr.download.williams_2018(). Verified against the PRIDE deposit used by the original publication: the two are now bit-identical with an identical missingness pattern.
Changed
Datasets (
pr.datasets) and Download (pr.download):williams_2018()gained azero_to_naparameter (defaultFalse), for consistency with sibling functions. It is mutually exclusive withfill_na. Note it governs zero semantics only and does not restore the summation defects above.Preprocessing (
pr.pp):normalize_median()now defaults to
log_space=Truerenamed the
batch_idparameter togroup_byrenamed the
zeros_to_naparameter tozero_to_na, for consistency with sibling functionsno longer accepts sparse
.Xinput
[0.1.1] - 2025-03-24
Added
Preprocessing (
pr.pp):summarize_modifications()for modification summarizationAnalysis (
pr.tl): ANOVA support indifferential_abundance()Visualization (
pr.pl):binary_heatmap(),box(),volcano(),peptides_on_sequence(),peptides_on_prot_sequence();print_statsparameter across multiple plot functionsDatasets (
pr.datasets):williams_2018()andkarayel_2020()download functionsUtilities (
pr.utils): Public API withis_proteodata(),check_proteodata(),is_log_transformed()Documentation: Sphinx documentation site; proteoform inference and protein-level analysis tutorials
Changed
Reader (
pr.read):diann()now supports version >=1.9.1 with automatic version dispatchPreprocessing (
pr.pp):impute_downshift()now supportsgroup_by;normalize_median()gainsmethodparameter;remove_contaminants()defaults toinplace=TrueValidation:
is_proteodata()now checks for NaN in ID columns, infinite values in.X/layers, and obs/var index sync
Fixed
volcano_plottype incompatibility and label displayn_cat1_per_cat2_histminimum bin width
[0.1.0] - 2025-01-29
Initial release of ProteoPy.
Added
Data import (
pr.read): Support for DIA-NN and generic long-format tablesAnnotation (
pr.ann): Functions to annotate samples (.obs) and variables (.var)Quality control (
pr.pp): Completeness filtering, CV calculation, contaminant removalPreprocessing (
pr.pp): Median normalization, downshift imputationDifferential abundance (
pr.tl): t-test, Welch’s test, ANOVA with multiple testing correctionProteoform inference (
pr.tl): COPF algorithm reimplementation for detecting functional proteoform groupsVisualization (
pr.pl): Volcano plots, abundance rank plots, intensity distributions, CV plots, correlation matrices, hierarchical clustering profilesDatasets (
pr.datasets): Built-in example datasets (Karayel 2020)